Nearly half of US inference openings are in the Bay Area
Open postings with inference or model serving in the title, grouped by US metro, as listed on the Metix AI Platform on September 21, 2026.
49.9%
176 of 353 distinct US inference postings are in the San Francisco Bay Area.
Bars are shares of all distinct postings; the dashed line is half
One metro holds nearly half. The next two hold another 23%.
Distinct US postings with inference or model serving in the title, by metro, as a share of all 353
Bar chart: San Francisco Bay Area 176 (49.9%), New York 43 (12.2%), Seattle 38 (10.8%); every other metro under 5%.
What it shows
The San Francisco Bay Area holds 176 of the 353 distinct US postings with inference or model serving in the title, 49.9%. New York (43) and Seattle (38) together hold another 23%. No other metro reaches 5%.
120 companies posted these roles, and five of them account for 118, about a third: NVIDIA (44), Capital One (21), Anthropic (20), Amazon Web Services (19), CoreWeave (14). The list mixes a chip maker, two cloud providers, a model lab, and a bank.
Method and limits
That 49.9% is a lower bound: the 37 postings that name no city are in the denominator. Among the 316 postings that do name a city, the Bay Area's share is 55.7%.
Worldwide, the Platform lists 585 open postings that match the same title filter, and 386 of them, about 66%, are in the United States. Both counts are before reposts are collapsed, so they compare with the 386 US postings read, not with the 353 distinct ones.
Source: Metix AI Platform, jobs, 2026-09-21.
Queriesworld-total.json · us-postings.json · us-postings-detail.json
{ "where": { "all": [ { "any": [ { "field": "title", "match": "inference" }, { "field": "title", "match": "model serving" }, { "field": "title", "match": "llm serving" } ] }, { "not": [ { "field": "title", "match": "causal" }, { "field": "title", "match": "statistical" } ] } ] }, "size": 1}
{ "where": { "all": [ { "any": [ { "field": "title", "match": "inference" }, { "field": "title", "match": "model serving" }, { "field": "title", "match": "llm serving" } ] }, { "not": [ { "field": "title", "match": "causal" }, { "field": "title", "match": "statistical" } ] }, { "field": "location.country", "eq": "United States" } ] }, "size": 10000}
{ "job_ids": "<the job_ids returned by us-postings.json, 100 per request>", "_source": [ "title", "company.name", "location.city", "location.state" ]}
Run it
Three ways in. Each says what it costs before you start.
1 API Credit buys 25 search results or 5 full records; $1 buys 30 API Credits.
Run it in your agent
96 to 110 API Credits$3.20 to $3.67May run past the 100 free API Credits
Your agent stops and asks before spending more than 150 API Credits.
Your agent follows the prompt step by step: it reads the rules, runs the counts, checks the definitions the prompt asks it to check, and writes the files and charts. Use an agent that can write files, such as Claude Code or Codex.
Set up onceKey, connection, and a free check. Skip this if your agent already reaches the Platform.
1Get a key
Create a key on the Metix AI Platform →New accounts get 100 API Credits once, valid for 30 days. Set the key in the shell you start your agent from, or add the line to ~/.zshrc or ~/.bashrc so every new terminal has it:
Shellexport METIX_KEY=metix_xxxxxxxx
2Connect your agent
Claude Code
Registers the Platform for every project. Start claude in any folder and the ten metix tools are there.
MCP setup guide →Shell: "${METIX_KEY:?run step 1 first}" && claude mcp add --scope user --transport http metix \ https://mira-api.metix.ai/mcp \ --header "Authorization: Bearer $METIX_KEY"Codex
Registers the same server. The key stays in your environment instead of the config file.
MCP setup guide →Shellcodex mcp add metix \ --url https://mira-api.metix.ai/mcp \ --bearer-token-env-var METIX_KEY
Skills
Four skills that teach any agent the Platform's endpoints and query rules, for agents without MCP. The installer starts with none ticked: press space on each, then enter.
Skills install guide →Shellnpx skills add MetixAI-Official/metix-skills
Other MCP
Point the client at this endpoint over streamable HTTP with both headers; without the Accept header the server answers 406. Older clients use /sse on the same host.
MCP setup guide →Endpoint and headershttps://mira-api.metix.ai/mcp Authorization: Bearer <your key> Accept: application/json, text/event-stream
3Check the setup
Ask this first. It reads your balance and the field list, runs no search, and costs nothing:
Prompt for your agentUse the Metix AI Platform to check my key status and read the contract; both are free. Then tell me my API Credit balance and which datasets I can query. Do not run any search.
4Paste the prompt
Start your agent in an empty folder, then paste. It writes its files there.
The question
Answer one question with the Metix AI Platform: where in the United States are companies hiring for model inference right now? Work only through the public Platform (REST at https://mira-api.metix.ai, the MCP server, or the metix-skills) with the key in METIX_KEY, and never print the key.
01Read before querying
Call GET /contract (free) and build every condition from querySpecByEntity.job. Read https://mira-api.metix.ai/docs/credits.md for prices: a search costs ceil(returned IDs / 25) API Credits, a detail read costs ceil(found records / 5), and a search that returns nothing is free.
02Population
Open job postings whose title matches any of "inference", "model serving", or "llm serving", minus titles that match "causal" or "statistical" (statistics roles, not model serving). Put the terms in one any node and the exclusions in one not node.
03Count first
With size 1, record the worldwide total and the total with location.country eq "United States". Each count costs 1 API Credit.
04Budget
Call GET /auth/key/status (free) now and again at the end, and report the difference as the cost. If the US total is under 1,000, read every US posting: search with size 10000, then POST /entity/v1/jobs/detail-by-id in batches of 100 with _source ["title", "company.name", "location.city", "location.state"]. Stop and ask before anything that would take the run past 150 API Credits.
05Audit
List the 40 most common titles and flag any that are not about serving or optimizing models in production. If more than 5% of postings are off-topic, add exclusions to step 2, rerun, and say what you changed.
06Clean
A posting with the same lowercased title, company, city, and state as an earlier one counts once. Report how many reposts this removed.
07Group by metro
Use six metros, San Francisco Bay Area, New York, Seattle, Washington and Baltimore, Boston, and Austin, each an explicit list of city and state pairs written to metros.json. Check the lists against the cities you actually read, so no large suburb falls outside its metro. A city on no list counts as "elsewhere in the US"; a missing city, or a state name in the city field, counts as "no city given". Keep both buckets.
08Outputs
Write data/metros.json (distinct postings and share by metro), data/companies.json (the five companies with the most distinct postings), and data/context.json (worldwide total, US total, records read, distinct postings), each with "unit": "jobs", the snapshot date, and the query file it came from. Keep the records in data/raw/ and never publish them.
09Chart
One horizontal bar chart of distinct postings by metro, sorted by count, with "elsewhere" and "no city given" last and in grey. Scale the bars to each metro's share of all distinct US postings, not to the longest bar, and mark 50%. Label every bar with its count and share. Title the chart with the finding, not the topic.
10Limits
Say what the numbers do not show: postings measure demand, not headcount; a role listed in several cities counts once per city; the index holds postings open on the snapshot day, and their posted dates are estimated.
What you get
The aggregate files and the chart, a note on what the audits found and what they changed, and the API Credits the run spent, read from the balance before and after.
A call that returns 402 insufficient_quota means the key works and the balance is empty.
Reproduce the numbers
96 API Credits$3.20Fits in the 100 free API Credits
A short standard-library Python script sends the committed queries, reads the 386 records the method needs, and writes the aggregate files this page is built from. It needs Python and METIX_KEY set in the shell (step 1 of the agent path); your agent can run these lines for you as well.
curl -fsSL https://platform.metix.ai/casebook/source/inference-roles-us-metros-2026.tar.gz | tar xz
cd inference-roles-us-metros-2026
: "${METIX_KEY:?set METIX_KEY first}" && python3 cases/inference-roles-us-metros-2026/fetch.pyWhat you get
data/*.json and data/receipt.json. Run git diff cases/inference-roles-us-metros-2026/data to see what moved: the numbers should match, apart from what changed in the data since the snapshot.
Adapt it
Costs what your version reads. Write your own ceiling into step 1 of the prompt.
The prompt is the case. Change the parts in this table and your agent answers your question instead, with the same checks and the same way of reporting cost.
| To change | Edit | For example |
|---|---|---|
| The role | The title terms and exclusions in step 2, and the audit rule in step 5 | "post-training" or "reinforcement learning", excluding "sales" |
| The country | The country in step 3 and the metro lists in step 7 | United Kingdom, with London, Cambridge, and Edinburgh |
| The grouping | Step 7 and the chart in step 9 | Group by company instead of by metro |
| The budget | The ceiling in step 4 | 20 API Credits: counts only, one count query per metro, no records |
What to ask before running it
When someone brings a looser version of this question ("where are the inference jobs?"), settle these first. Each one changes the query or the cost:
- Which roles, in title words, and what should be excluded?
- Which geography, and at what level: country, state, or metro?
- Which postings: everything open today, or only those posted in the last 30 or 90 days?
- What to count: postings, distinct postings, or companies?
- How many API Credits the run may spend.
Answer one question with the Metix AI Platform: where in the United States are companies hiring for model inference right now? Work only through the public Platform (REST at https://mira-api.metix.ai, the MCP server, or the metix-skills) with the key in METIX_KEY, and never print the key. 1. Read before querying. Call GET /contract (free) and build every condition from querySpecByEntity.job. Read https://mira-api.metix.ai/docs/credits.md for prices: a search costs ceil(returned IDs / 25) API Credits, a detail read costs ceil(found records / 5), and a search that returns nothing is free. 2. Population. Open job postings whose title matches any of "inference", "model serving", or "llm serving", minus titles that match "causal" or "statistical" (statistics roles, not model serving). Put the terms in one any node and the exclusions in one not node. 3. Count first. With size 1, record the worldwide total and the total with location.country eq "United States". Each count costs 1 API Credit. 4. Budget. Call GET /auth/key/status (free) now and again at the end, and report the difference as the cost. If the US total is under 1,000, read every US posting: search with size 10000, then POST /entity/v1/jobs/detail-by-id in batches of 100 with _source ["title", "company.name", "location.city", "location.state"]. Stop and ask before anything that would take the run past 150 API Credits. 5. Audit. List the 40 most common titles and flag any that are not about serving or optimizing models in production. If more than 5% of postings are off-topic, add exclusions to step 2, rerun, and say what you changed. 6. Clean. A posting with the same lowercased title, company, city, and state as an earlier one counts once. Report how many reposts this removed. 7. Group by metro. Use six metros, San Francisco Bay Area, New York, Seattle, Washington and Baltimore, Boston, and Austin, each an explicit list of city and state pairs written to metros.json. Check the lists against the cities you actually read, so no large suburb falls outside its metro. A city on no list counts as "elsewhere in the US"; a missing city, or a state name in the city field, counts as "no city given". Keep both buckets. 8. Outputs. Write data/metros.json (distinct postings and share by metro), data/companies.json (the five companies with the most distinct postings), and data/context.json (worldwide total, US total, records read, distinct postings), each with "unit": "jobs", the snapshot date, and the query file it came from. Keep the records in data/raw/ and never publish them. 9. Chart. One horizontal bar chart of distinct postings by metro, sorted by count, with "elsewhere" and "no city given" last and in grey. Scale the bars to each metro's share of all distinct US postings, not to the longest bar, and mark 50%. Label every bar with its count and share. Title the chart with the finding, not the topic. 10. Limits. Say what the numbers do not show: postings measure demand, not headcount; a role listed in several cities counts once per city; the index holds postings open on the snapshot day, and their posted dates are estimated.
Method and limits
How the population was defined, counted, and checked, and what the numbers cannot show.
Population
Open job postings on the Metix AI Platform on September 21, 2026, whose title matches "inference", "model serving", or "llm serving", excluding titles that match "causal" or "statistical". The exclusions remove statistics roles such as causal inference, which are not about serving models.
Counting
A distinct posting is one title, company, city, and state. The replay read 386 US postings and collapsed 33 same-city reposts, leaving 353. A role listed in three cities counts once in each.
Metros
Each metro is a list of cities by state, in metros.json. A city on no list counts as elsewhere in the US. A posting with no city, or with a state name in the city field, counts as no city given.
How it was made
An AI agent built this case through the public REST API. Before writing the replay it read every matching US posting once. That read showed that 40 of 426 were statistics roles and that reposts inflate city counts. The queries it settled on are in queries/, and fetch.py replays them without an agent.
Limits
Postings measure demand, not headcount. A company that lists one role in several metros counts once in each. The 37 postings with no city stay in the denominator, so every metro's share is a lower bound. The index holds postings open on the snapshot day, and their posted dates are estimated, so this is one day, not a trend. Inference work under other titles ("ML performance engineer", for example) is not counted, so the totals are a lower bound too.
The last reproduction
What reproducing this case cost the last time the script ran, read from the Platform's own balance before and after.
- Ran on
- 2026-09-21
- Calls
- 7
- Search results
- 388
- Records read
- 386
- API Credits
- 96
Search results are IDs returned by searches, one per count query and one per match on a full search. Records are postings or profiles read in full: 386 here.
Making this case cost about 206 API Credits more: the agent's audits, trial queries, and earlier runs that the published replay replaced. You do not pay that again.