AKVCAK Venture Corp

Report · llm-inference · updated 2026-09-24

LLM inference platforms

A timely slice of AI infrastructure: platforms that serve tokens and models to other software. Mix of neoclouds, GPU serverless, and the open-source serving engine many of them wrap. Only public sources are cited.

DIRECTEDANCHOREDFORMINGOPERATIONALGroqFireworks AITogether AICerebras InferenceModalReplicateBasetenvLLMHugging Face InferenceCloudflare Workers AIPROVEN EXECUTION →VALIDATED DIRECTION →

Placements

10 of 10 entities have a Jev placement. Axes: Proven Execution × Validated Direction.

Data table alternative to the LLM inference platforms quadrant chart
EntityStatusProven ExecutionValidated DirectionuXuYRegionP(region)
Groqscored0.4850.4480.3590.382DIRECTEDOPERATIONAL:0.27 ANCHORED:0.25 FORMING:0.06 DIRECTED:0.42
Fireworks AIscored0.3670.2630.3510.349OPERATIONALANCHORED:0.03 FORMING:0.14 DIRECTED:0.06 OPERATIONAL:0.77
Together AIscored0.3570.2100.3180.253OPERATIONALFORMING:0.14 OPERATIONAL:0.77 DIRECTED:0.06 ANCHORED:0.03
Cerebras Inferencescored0.3130.3230.3580.369OPERATIONALANCHORED:0.04 DIRECTED:0.01 FORMING:0.01 OPERATIONAL:0.94
Modalscored0.4680.2250.3250.269OPERATIONALFORMING:0.23 ANCHORED:0.03 OPERATIONAL:0.61 DIRECTED:0.13
Replicatescored0.3880.2150.4110.344OPERATIONALANCHORED:0.10 OPERATIONAL:0.81 DIRECTED:0.03 FORMING:0.06
Basetenscored0.2550.2350.3530.343OPERATIONALFORMING:0.18 OPERATIONAL:0.74 DIRECTED:0.06 ANCHORED:0.02
vLLMscored0.3950.2300.3940.347OPERATIONALANCHORED:0.03 DIRECTED:0.02 OPERATIONAL:0.81 FORMING:0.14
Hugging Face Inferencescored0.2800.2020.3650.328OPERATIONALANCHORED:0.02 DIRECTED:0.07 OPERATIONAL:0.60 FORMING:0.31
Cloudflare Workers AIscored0.2320.2470.3500.359OPERATIONALOPERATIONAL:0.56 ANCHORED:0.02 FORMING:0.33 DIRECTED:0.08

Entities

  • Groq

    LPU-based inference cloud (GroqCloud) and silicon now in the NVIDIA stack.

    2 evidence · scored · DIRECTED

  • Fireworks AI

    Training and inference platform for open and specialized models.

    1 evidence · scored · OPERATIONAL

  • Together AI

    Open-model inference, fine-tuning, and GPU cloud.

    2 evidence · scored · OPERATIONAL

  • Cerebras Inference

    Wafer-scale hosted inference with an OpenAI-compatible Chat Completions API.

    1 evidence · scored · OPERATIONAL

  • Modal

    Programmable serverless GPU compute used for training and serving.

    2 evidence · scored · OPERATIONAL

  • Replicate

    API for running and hosting public and custom models, often billed by runtime.

    1 evidence · scored · OPERATIONAL

  • Baseten

    Model APIs plus dedicated deployments for custom serving.

    1 evidence · scored · OPERATIONAL

  • vLLM

    Open-source high-throughput LLM serving engine used under many hosted stacks.

    1 evidence · scored · OPERATIONAL

  • Hugging Face Inference

    Hosted inference and serverless endpoints on the Hugging Face Hub.

    1 evidence · scored · OPERATIONAL

  • Cloudflare Workers AI

    Inference on Cloudflare’s edge network via Workers bindings.

    1 evidence · scored · OPERATIONAL

Payment does not affect scores. Submit evidence or read the rubric.