Scoring Architecture

How JobsVsAI scores work.

A calm, evidence-led breakdown of what our scores measure, how they are calculated from task-level data, and where the boundaries of certainty lie.

AI Exposure
0–100

How much of this occupation’s actual task mix current AI can meaningfully act on. It is a statement about the work, not about your job.

Replacement Risk
0–100

How structurally vulnerable the human role is to reduced labour demand after physical reality, human dependency, accountability, regulation and adoption are accounted for.

These two are not the same number, and the gap is the point.

A surgeon has tasks current AI can assist with, and a replacement risk that stays low because the work requires physical presence, carries severe consequences, and someone has to be accountable for the outcome. A collapsing of exposure into replacement would get that backwards. Both scores are indices, not probabilities.

7-Step Measurement Process

From raw tasks to calibrated career intelligence

How occupational data flows through our capability and constraint models.

  1. 01
    Decompose occupations into individual tasks

    We break each role down into official O*NET 30.3 task statements, weighted by importance and frequency ratings.

  2. 02
    Map tasks to capability requirements

    Each task is evaluated across 15 AI capability dimensions (e.g. natural language synthesis, pattern recognition, spatial navigation).

  3. 03
    Measure current commercial AI capability

    We evaluate verified frontier and commercial AI systems against required task levels using our Capability Index.

  4. 04
    Estimate task automation feasibility

    Capability is modified by real-world friction: environmental variability, consequence severity, and regulatory burden.

  5. 05
    Apply human and environmental barriers

    We evaluate physical dependency (manual dexterity, mobility) and human dependency (interpersonal trust, ethics, accountability).

  6. 06
    Calculate AI Exposure and Replacement Risk separately

    Task exposure reflects capability overlap; Replacement Risk incorporates structural adoption pressure and labour resilience.

  7. 07
    Publish only after rigorous evidence gating

    Occupations publish as verified only when weighted task coverage exceeds 80% and confidence passes validation thresholds.

Task level

Three questions, kept separate

Occupations are decomposed into their O*NET tasks, and each task is assessed three ways. Collapsing these into one number is the most common way AI-risk estimates go wrong.

AI Capability Fit

Does current commercially deployable AI have the capabilities this task requires? Computed as a weighted geometric mean across fifteen capability dimensions, so a critical weakness cannot be averaged away by strength elsewhere.

Automation Feasibility

Even where AI is capable, can this task actually be automated in its real working environment? Physical presence, variability, regulation, accountability and consequence severity all reduce it.

Augmentation Potential

Could AI substantially help a human do this task, even where full automation is unrealistic? Reported per task. There is deliberately no occupation-level augmentation headline: that number has not been validated.

Occupation level

Six factors, and two of them are provisional

Weights are fixed in versioned scoring configuration. Two factors — adoption pressure (weight 0.15) and labour-market resilience resistance (weight 0.10), together 25% of Replacement Risk weighting — rest on models we consider provisional, and every occupation page reports how sensitive its score is to them. Provisional means estimated from structural proxies rather than measured directly, and not yet through the validation the other four factors have had.

35%

Task automation exposure

The importance-and-frequency-weighted average of how feasible it is to automate each task in the occupation — not how capable AI is in the abstract.

10%

AI capability proximity

How close current commercially deployable AI is to the capabilities the work actually requires, before real-world constraints are applied.

15%

Human dependency resistance

Trust, judgement, accountability and relationship work, derived from O*NET evidence about the occupation rather than assumed by category.

15%

Physical dependency resistance

Physical presence, manipulation and mobility requirements, reconstructed directly from O*NET work-context evidence.

15%

Adoption pressure Provisional

How readily employers reorganise this work around AI. A provisional structural model, disclosed transparently as such.

10%

Labour-market resilience resistance Provisional

Demand and sector conditions. Also provisional, and disclosed rather than quietly folded into the headline number.

The bottleneck principle

Strength in one capability does not cancel weakness in another

If a task is 60% fine physical manipulation and 40% language, excellent language ability does not make the task automatable. Capability Fit uses a weighted geometric mean and applies an explicit cap when a critical requirement is unmet, so a genuine bottleneck survives into the final score instead of being averaged out of it.

The Frontier AI Capability Index

JobsVsAI maintains its own index of what commercially deployable AI can currently do across fifteen capability dimensions — from language comprehension to mobility in the physical world. It is a synthesis over independent evaluations, vendor evaluations, academic research and documented deployments, not an average of benchmark scores.

A separate technical-frontier track exists and is deliberately empty: we have not seen evidence sufficient to populate it responsibly.

Coverage and confidence

We would rather publish nothing than publish a guess

An occupation is only scored when at least 70% of its weighted task evidence is usable. Below that it stays unpublished — no default values, no borrowed category averages, no filling gaps with what similar occupations look like.

Confidence is reported as a number out of 100, not a High/Medium/Low badge, and combines weighted coverage, mapping quality, capability-evidence quality, source completeness and proxy confidence.

  1. 01
    Map tasks to capability requirements

    What the work requires — assessed independently of what AI can currently do, so capability updates do not require remapping.

  2. 02
    Apply the current capability index

    Producing Capability Fit, Automation Feasibility and Augmentation Potential per task.

  3. 03
    Aggregate with structural constraints

    Weighted by task importance and frequency, then adjusted for real-world constraints.

  4. 04
    Gate, version and persist

    Coverage and confidence gates decide publishability. Every score keeps its inputs, weights and formula versions.

Two score classes

Verified analyses and preliminary estimates

A verified JobsVsAI analysis has cleared every gate above: at least 80% weighted task coverage, a confidence threshold, mapping completeness, and a review of the factors carrying provisional models. 507 occupations currently qualify.

A preliminary estimate has not. It exists because returning nothing for an occupation we know something real about serves nobody — but it is never presented as a verified score, never enters our rankings, and is labelled before any number appears.

How estimates are made. Every estimate is deterministic and uses only data already imported. Nothing is inferred from an occupation’s title.

  • Complete task evidence. The same calculation used for verified occupations, over full task coverage. Withheld from the verified cohort by a review gate rather than by missing evidence.
  • Partial task evidence. The same calculation over the task evidence available so far, which covers only part of the work.
  • Related-occupation estimate. No task evidence exists, so the figure is drawn from fully analysed occupations that O*NET itself identifies as closely related, weighted by how close that relationship is. These are always shown as a range.

How accurate are they? We test the related-occupation method by taking each of the 507 verified occupations, hiding its own evidence, and estimating it from its relatives alone. Half the estimates land within 3.6 points of the verified AI Exposure score and 2.8 points of Replacement Risk; nine in ten land within 10.2 and 7.7 points respectively. Roughly 78% land in the same risk band. That is useful, and it is not the same as measured.

Limits. An estimate carries no task-level breakdown, no career transitions and no action plan, because those need validated task evidence and we will not generate guidance we cannot support. An estimate becomes a verified analysis when the underlying evidence clears the gates — usually when its tasks are mapped, or when a provisional factor model is validated. Nothing about the estimate is carried over; the verified score is calculated from scratch.

  1. 01
    Never from a title

    An occupation’s name is not evidence. Every estimate traces to imported task data or to named, fully analysed related occupations.

  2. 02
    Labelled before the number

    The preliminary status and its confidence appear above the scores, not in a footnote beneath them.

  3. 03
    Ranges where precision is not earned

    Where the evidence supports a span rather than a point, we show the span.

  4. 04
    Excluded from rankings

    Our highest and lowest replacement-risk lists contain verified occupations only, so an estimate can never distort them.

Versioning

Every score can be rebuilt

Scores are immutable snapshots. Each records the frontier index version, structural proxy model, occupation and task formula versions, capability taxonomy, mapping rubric and evidence policy that produced it. The same inputs and versions always reproduce the same number.

When frontier AI capability changes, we update the capability index and recalculate — without rebuilding the occupational knowledge base underneath it.

Occupation data
O*NET 30.3, on source release
Capability index
On material capability shifts
Structural proxies
Direct O*NET evidence
Score snapshots
Immutable, versioned, promoted in runs
Read Technical Methodology →View Methodology Changelog →
What these scores are not.

They are decision-support indices, not probabilities that any individual will lose their job. Job content varies by employer, seniority and country. Adoption pressure and labour-market resilience remain provisional models, and occupations whose scores depend heavily on them are held back from publication rather than shipped with a caveat. Where our evidence is thin, the occupation does not appear at all.

Source attribution

Occupational data from O*NET 30.3 by the U.S. Department of Labor, Employment and Training Administration. Used under CC BY 4.0. O*NET® is a trademark of USDOL/ETA. JobsVsAI scores, capability taxonomy and structural models are our own interpretation and are not endorsed by USDOL/ETA.