AI infrastructure company Infinity raised at a $100M valuation — a bet that inference optimization is the next chapter of the AI stack.
Infinity, a startup building inference-optimization tooling for large language models, closed a $15 million round at a $100 million valuation. The round was led by Touring Capital, with participation from Principal VC and individual researchers from OpenAI and Anthropic.
Infinity's core product is a runtime layer that sits between model calls and the underlying compute, applying speculative decoding, KV-cache reuse, and adaptive batching to squeeze latency and cost out of production LLM workloads.
The check size is modest — but the signal is loud. When researchers from the labs whose models you're optimizing are personally investing, it tells enterprises the approach is real. This is the classic pre-scale validation moment.
Inference optimization is having a moment. Together AI, Anyscale, Baseten, and Modal all target adjacent slices of this problem. Infinity's bet is that a purpose-built runtime beats a general-purpose platform on the specific problem of getting more out of the same GPU.
Source: TechCrunch