Report · llm-inference · updated 2026-09-24
LLM inference platforms
A timely slice of AI infrastructure: platforms that serve tokens and models to other software. Mix of neoclouds, GPU serverless, and the open-source serving engine many of them wrap. Only public sources are cited.
Placements
10 of 10 entities have a Jev placement. Axes: Proven Execution × Validated Direction.
| Entity | Status | Proven Execution | Validated Direction | uX | uY | Region | P(region) |
|---|---|---|---|---|---|---|---|
| Groq | scored | 0.485 | 0.448 | 0.359 | 0.382 | DIRECTED | OPERATIONAL:0.27 ANCHORED:0.25 FORMING:0.06 DIRECTED:0.42 |
| Fireworks AI | scored | 0.367 | 0.263 | 0.351 | 0.349 | OPERATIONAL | ANCHORED:0.03 FORMING:0.14 DIRECTED:0.06 OPERATIONAL:0.77 |
| Together AI | scored | 0.357 | 0.210 | 0.318 | 0.253 | OPERATIONAL | FORMING:0.14 OPERATIONAL:0.77 DIRECTED:0.06 ANCHORED:0.03 |
| Cerebras Inference | scored | 0.313 | 0.323 | 0.358 | 0.369 | OPERATIONAL | ANCHORED:0.04 DIRECTED:0.01 FORMING:0.01 OPERATIONAL:0.94 |
| Modal | scored | 0.468 | 0.225 | 0.325 | 0.269 | OPERATIONAL | FORMING:0.23 ANCHORED:0.03 OPERATIONAL:0.61 DIRECTED:0.13 |
| Replicate | scored | 0.388 | 0.215 | 0.411 | 0.344 | OPERATIONAL | ANCHORED:0.10 OPERATIONAL:0.81 DIRECTED:0.03 FORMING:0.06 |
| Baseten | scored | 0.255 | 0.235 | 0.353 | 0.343 | OPERATIONAL | FORMING:0.18 OPERATIONAL:0.74 DIRECTED:0.06 ANCHORED:0.02 |
| vLLM | scored | 0.395 | 0.230 | 0.394 | 0.347 | OPERATIONAL | ANCHORED:0.03 DIRECTED:0.02 OPERATIONAL:0.81 FORMING:0.14 |
| Hugging Face Inference | scored | 0.280 | 0.202 | 0.365 | 0.328 | OPERATIONAL | ANCHORED:0.02 DIRECTED:0.07 OPERATIONAL:0.60 FORMING:0.31 |
| Cloudflare Workers AI | scored | 0.232 | 0.247 | 0.350 | 0.359 | OPERATIONAL | OPERATIONAL:0.56 ANCHORED:0.02 FORMING:0.33 DIRECTED:0.08 |
Entities
- Groq
LPU-based inference cloud (GroqCloud) and silicon now in the NVIDIA stack.
2 evidence · scored · DIRECTED
- Fireworks AI
Training and inference platform for open and specialized models.
1 evidence · scored · OPERATIONAL
- Together AI
Open-model inference, fine-tuning, and GPU cloud.
2 evidence · scored · OPERATIONAL
- Cerebras Inference
Wafer-scale hosted inference with an OpenAI-compatible Chat Completions API.
1 evidence · scored · OPERATIONAL
- Modal
Programmable serverless GPU compute used for training and serving.
2 evidence · scored · OPERATIONAL
- Replicate
API for running and hosting public and custom models, often billed by runtime.
1 evidence · scored · OPERATIONAL
- Baseten
Model APIs plus dedicated deployments for custom serving.
1 evidence · scored · OPERATIONAL
- vLLM
Open-source high-throughput LLM serving engine used under many hosted stacks.
1 evidence · scored · OPERATIONAL
- Hugging Face Inference
Hosted inference and serverless endpoints on the Hugging Face Hub.
1 evidence · scored · OPERATIONAL
- Cloudflare Workers AI
Inference on Cloudflare’s edge network via Workers bindings.
1 evidence · scored · OPERATIONAL
Payment does not affect scores. Submit evidence or read the rubric.