Card 03Jobs

Nearly half of US inference openings are in the Bay Area

Open postings with inference or model serving in the title, grouped by US metro, as listed on the Metix AI Platform on September 21, 2026.

49.9%

176 of 353 distinct US inference postings are in the San Francisco Bay Area.

  • San Francisco Bay Area49.9%
  • New York12.2%
  • Seattle10.8%
  • Other US metros and no city27.2%

Bars are shares of all distinct postings; the dashed line is half

One metro holds nearly half. The next two hold another 23%.

Distinct US postings with inference or model serving in the title, by metro, as a share of all 353

Bar chart: San Francisco Bay Area 176 (49.9%), New York 43 (12.2%), Seattle 38 (10.8%); every other metro under 5%.

  1. San Francisco Bay Area176
  2. New York43
  3. Seattle38
  4. Washington and Baltimore16
  5. Boston11
  6. Austin9
  7. Elsewhere in the US23
  8. No city given37

What it shows

The San Francisco Bay Area holds 176 of the 353 distinct US postings with inference or model serving in the title, 49.9%. New York (43) and Seattle (38) together hold another 23%. No other metro reaches 5%.

120 companies posted these roles, and five of them account for 118, about a third: NVIDIA (44), Capital One (21), Anthropic (20), Amazon Web Services (19), CoreWeave (14). The list mixes a chip maker, two cloud providers, a model lab, and a bank.

Method and limits

That 49.9% is a lower bound: the 37 postings that name no city are in the denominator. Among the 316 postings that do name a city, the Bay Area's share is 55.7%.

Worldwide, the Platform lists 585 open postings that match the same title filter, and 386 of them, about 66%, are in the United States. Both counts are before reposts are collapsed, so they compare with the 386 US postings read, not with the 353 distinct ones.

Source: Metix AI Platform, jobs, 2026-09-21.

Queriesworld-total.json · us-postings.json · us-postings-detail.json
POST /v1/jobs/query · queries/world-total.json
{  "where": {    "all": [      {        "any": [          {            "field": "title",            "match": "inference"          },          {            "field": "title",            "match": "model serving"          },          {            "field": "title",            "match": "llm serving"          }        ]      },      {        "not": [          {            "field": "title",            "match": "causal"          },          {            "field": "title",            "match": "statistical"          }        ]      }    ]  },  "size": 1}
POST /v1/jobs/query · queries/us-postings.json
{  "where": {    "all": [      {        "any": [          {            "field": "title",            "match": "inference"          },          {            "field": "title",            "match": "model serving"          },          {            "field": "title",            "match": "llm serving"          }        ]      },      {        "not": [          {            "field": "title",            "match": "causal"          },          {            "field": "title",            "match": "statistical"          }        ]      },      {        "field": "location.country",        "eq": "United States"      }    ]  },  "size": 10000}
POST /entity/v1/jobs/detail-by-id · queries/us-postings-detail.json
{  "job_ids": "<the job_ids returned by us-postings.json, 100 per request>",  "_source": [    "title",    "company.name",    "location.city",    "location.state"  ]}

Run it

Three ways in. Each says what it costs before you start.

1 API Credit buys 25 search results or 5 full records; $1 buys 30 API Credits.

Run it in your agent

96 to 110 API Credits$3.20 to $3.67May run past the 100 free API Credits

Your agent stops and asks before spending more than 150 API Credits.

Your agent follows the prompt step by step: it reads the rules, runs the counts, checks the definitions the prompt asks it to check, and writes the files and charts. Use an agent that can write files, such as Claude Code or Codex.

Set up onceKey, connection, and a free check. Skip this if your agent already reaches the Platform.
  1. 1Get a key

    Create a key on the Metix AI Platform →

    New accounts get 100 API Credits once, valid for 30 days. Set the key in the shell you start your agent from, or add the line to ~/.zshrc or ~/.bashrc so every new terminal has it:

    Shell
    export METIX_KEY=metix_xxxxxxxx
  2. 2Connect your agent

    Claude Code

    Registers the Platform for every project. Start claude in any folder and the ten metix tools are there.

    Shell
    : "${METIX_KEY:?run step 1 first}" &&
    claude mcp add --scope user --transport http metix \
      https://mira-api.metix.ai/mcp \
      --header "Authorization: Bearer $METIX_KEY"
    MCP setup guide →

    Codex

    Registers the same server. The key stays in your environment instead of the config file.

    Shell
    codex mcp add metix \
      --url https://mira-api.metix.ai/mcp \
      --bearer-token-env-var METIX_KEY
    MCP setup guide →

    Skills

    Four skills that teach any agent the Platform's endpoints and query rules, for agents without MCP. The installer starts with none ticked: press space on each, then enter.

    Shell
    npx skills add MetixAI-Official/metix-skills
    Skills install guide →

    Other MCP

    Point the client at this endpoint over streamable HTTP with both headers; without the Accept header the server answers 406. Older clients use /sse on the same host.

    Endpoint and headers
    https://mira-api.metix.ai/mcp
    Authorization: Bearer <your key>
    Accept: application/json, text/event-stream
    MCP setup guide →
  3. 3Check the setup

    Ask this first. It reads your balance and the field list, runs no search, and costs nothing:

    Prompt for your agent
    Use the Metix AI Platform to check my key status and read the contract; both are free. Then tell me my API Credit balance and which datasets I can query. Do not run any search.

4Paste the prompt

Start your agent in an empty folder, then paste. It writes its files there.

The question

Answer one question with the Metix AI Platform: where in the United States are companies hiring for model inference right now? Work only through the public Platform (REST at https://mira-api.metix.ai, the MCP server, or the metix-skills) with the key in METIX_KEY, and never print the key.

  1. 01Read before querying

    Call GET /contract (free) and build every condition from querySpecByEntity.job. Read https://mira-api.metix.ai/docs/credits.md for prices: a search costs ceil(returned IDs / 25) API Credits, a detail read costs ceil(found records / 5), and a search that returns nothing is free.

  2. 02Population

    Open job postings whose title matches any of "inference", "model serving", or "llm serving", minus titles that match "causal" or "statistical" (statistics roles, not model serving). Put the terms in one any node and the exclusions in one not node.

  3. 03Count first

    With size 1, record the worldwide total and the total with location.country eq "United States". Each count costs 1 API Credit.

  4. 04Budget

    Call GET /auth/key/status (free) now and again at the end, and report the difference as the cost. If the US total is under 1,000, read every US posting: search with size 10000, then POST /entity/v1/jobs/detail-by-id in batches of 100 with _source ["title", "company.name", "location.city", "location.state"]. Stop and ask before anything that would take the run past 150 API Credits.

  5. 05Audit

    List the 40 most common titles and flag any that are not about serving or optimizing models in production. If more than 5% of postings are off-topic, add exclusions to step 2, rerun, and say what you changed.

  6. 06Clean

    A posting with the same lowercased title, company, city, and state as an earlier one counts once. Report how many reposts this removed.

  7. 07Group by metro

    Use six metros, San Francisco Bay Area, New York, Seattle, Washington and Baltimore, Boston, and Austin, each an explicit list of city and state pairs written to metros.json. Check the lists against the cities you actually read, so no large suburb falls outside its metro. A city on no list counts as "elsewhere in the US"; a missing city, or a state name in the city field, counts as "no city given". Keep both buckets.

  8. 08Outputs

    Write data/metros.json (distinct postings and share by metro), data/companies.json (the five companies with the most distinct postings), and data/context.json (worldwide total, US total, records read, distinct postings), each with "unit": "jobs", the snapshot date, and the query file it came from. Keep the records in data/raw/ and never publish them.

  9. 09Chart

    One horizontal bar chart of distinct postings by metro, sorted by count, with "elsewhere" and "no city given" last and in grey. Scale the bars to each metro's share of all distinct US postings, not to the longest bar, and mark 50%. Label every bar with its count and share. Title the chart with the finding, not the topic.

  10. 10Limits

    Say what the numbers do not show: postings measure demand, not headcount; a role listed in several cities counts once per city; the index holds postings open on the snapshot day, and their posted dates are estimated.

What you get

The aggregate files and the chart, a note on what the audits found and what they changed, and the API Credits the run spent, read from the balance before and after.

A call that returns 402 insufficient_quota means the key works and the balance is empty.

Reproduce the numbers

96 API Credits$3.20Fits in the 100 free API Credits

A short standard-library Python script sends the committed queries, reads the 386 records the method needs, and writes the aggregate files this page is built from. It needs Python and METIX_KEY set in the shell (step 1 of the agent path); your agent can run these lines for you as well.

Shell
curl -fsSL https://platform.metix.ai/casebook/source/inference-roles-us-metros-2026.tar.gz | tar xz
cd inference-roles-us-metros-2026
: "${METIX_KEY:?set METIX_KEY first}" && python3 cases/inference-roles-us-metros-2026/fetch.py

What you get

data/*.json and data/receipt.json. Run git diff cases/inference-roles-us-metros-2026/data to see what moved: the numbers should match, apart from what changed in the data since the snapshot.

Adapt it

Costs what your version reads. Write your own ceiling into step 1 of the prompt.

The prompt is the case. Change the parts in this table and your agent answers your question instead, with the same checks and the same way of reporting cost.

To change Edit For example
The role The title terms and exclusions in step 2, and the audit rule in step 5 "post-training" or "reinforcement learning", excluding "sales"
The country The country in step 3 and the metro lists in step 7 United Kingdom, with London, Cambridge, and Edinburgh
The grouping Step 7 and the chart in step 9 Group by company instead of by metro
The budget The ceiling in step 4 20 API Credits: counts only, one count query per metro, no records

What to ask before running it

When someone brings a looser version of this question ("where are the inference jobs?"), settle these first. Each one changes the query or the cost:

  1. Which roles, in title words, and what should be excluded?
  2. Which geography, and at what level: country, state, or metro?
  3. Which postings: everything open today, or only those posted in the last 30 or 90 days?
  4. What to count: postings, distinct postings, or companies?
  5. How many API Credits the run may spend.

Method and limits

How the population was defined, counted, and checked, and what the numbers cannot show.

Population

Open job postings on the Metix AI Platform on September 21, 2026, whose title matches "inference", "model serving", or "llm serving", excluding titles that match "causal" or "statistical". The exclusions remove statistics roles such as causal inference, which are not about serving models.

Counting

A distinct posting is one title, company, city, and state. The replay read 386 US postings and collapsed 33 same-city reposts, leaving 353. A role listed in three cities counts once in each.

Metros

Each metro is a list of cities by state, in metros.json. A city on no list counts as elsewhere in the US. A posting with no city, or with a state name in the city field, counts as no city given.

How it was made

An AI agent built this case through the public REST API. Before writing the replay it read every matching US posting once. That read showed that 40 of 426 were statistics roles and that reposts inflate city counts. The queries it settled on are in queries/, and fetch.py replays them without an agent.

Limits

Postings measure demand, not headcount. A company that lists one role in several metros counts once in each. The 37 postings with no city stay in the denominator, so every metro's share is a lower bound. The index holds postings open on the snapshot day, and their posted dates are estimated, so this is one day, not a trend. Inference work under other titles ("ML performance engineer", for example) is not counted, so the totals are a lower bound too.

The last reproduction

What reproducing this case cost the last time the script ran, read from the Platform's own balance before and after.

Ran on
2026-09-21
Calls
7
Search results
388
Records read
386
API Credits
96

Search results are IDs returned by searches, one per count query and one per match on a full search. Records are postings or profiles read in full: 386 here.

Making this case cost about 206 API Credits more: the agent's audits, trial queries, and earlier runs that the published replay replaced. You do not pay that again.