HomeAI NewsKog's software unlocks faster inference on GPUs enterprises already own

Kog’s software unlocks faster inference on GPUs enterprises already own

The French startup’s software-only approach aims for 30x faster LLM decoding on conventional data center GPUs.

French startup Kog has unveiled a software approach that dramatically accelerates inference on conventional GPUs from AMD and NVIDIA. The company’s tech preview achieved 3,000 tokens per second on a small model, and it now aims to deliver the same speed for large language models.

Kog’s demo used AMD MI300X and NVIDIA H200 GPUs, and the company open-sourced the Laneformer 2B model. The startup earned 200 business leads after its Hacker News debut, with software engineering as the first target use case. CEO Gaël Delalleau, whose first startup was Stribe, compares Kog’s work to Stanford’s Hazy Research.

For builders, the appeal is avoiding expensive GPU upgrades and time-consuming model fine-tuning. Kog’s software targets enterprise workflows like AI coding assistants, where wait times stall professional output. Kog also partners with low-code platforms that generate games and apps, where faster inference translates directly into more revenue.

Kog faces a major test in scaling its methods to large language models. The company’s demo used a 2-billion-parameter model, and its goal is a 30x speedup for LLMs on standard hardware. The company must prove that its software delivers on that promise when larger models run in production.

What matters

  • Kog’s tech preview hit 3,000 tokens per second on standard AMD and NVIDIA data center GPUs.
  • Software-only optimization can unlock faster inference without replacing existing GPU infrastructure.
  • Watch whether Kog scales its approach from a 2-billion-parameter model to large LLMs.

Why it matters

Watch whether Kog scales its approach from a 2-billion-parameter model to large LLMs.

This GenAI News article was prepared in original wording using reporting and materials published by TechCrunch AI. Source reference: https://techcrunch.com/2026/08/14/kog-is-going-deeper-to-squeeze-more-inference-out-of-gpus/.

Drafted by the GenAI News review pipeline.

latest articles

explore more