
Unlocking higher inference speeds on existing graphics processing units
French startup Kog is using low-level software optimization to extract faster inference from standard datacenter hardware. By applying principles from solid-state physics and reverse engineering, the company aims to prove that conventional chips still have untapped potential.
Published by Jin · 2 min read · 15 AUG 2026
The race for faster artificial intelligence inference has driven significant market interest in purpose-built hardware. However, a French startup named Kog is taking a different approach by focusing on software optimization for the conventional datacenter graphics processing units that enterprises already own, such as the AMD MI300X and Nvidia H200.
Targeting existing hardware
Kog gained attention with a technical preview demonstrating extremely fast single-request decoding on standard hardware. With inference speed and operational costs acting as critical bottlenecks, the promise of unlocking new capabilities through software alone generated substantial business interest following its initial public debut in May.
The startup's initial focus centers on software engineering use cases. Professionals utilizing tools like Claude Code frequently experience delays, and companies often charge premiums for faster execution modes. Kog hopes to capture customers put off by these bottlenecks, including design partners generating games and apps via text prompts.
The engineering approach
Led by CEO Gaël Delalleau, Kog relies on a deep-level focus on graphics processing unit acceleration. Delalleau studied solid-state physics at France's École Polytechnique and spent time in offensive cybersecurity. This background shaped an engineering philosophy centered on understanding the fundamental laws of hardware and reverse-engineering systems down to assembly language and binary code to achieve unintended goals.
This hands-on methodology requires significant time and effort. For every new hardware architecture, the eleven-person team dedicates weeks or months to conduct detailed engineering research. This limits the initial number of supported chips, though the startup ultimately hopes to feed its methodology into automated pipelines.
Scaling to larger models
While the company's initial demonstration achieved 3,000 tokens per second, it relied on a purpose-built small model known as Laneformer 2B with approximately two billion parameters. Prospective customers, however, are generally not prepared to fine-tune small models for production environments.
Consequently, Kog has shifted its development efforts toward accelerating larger models. The startup is backed by local investors including Varsity VC, Bpifrance, and Scaleway, and aims to demonstrate implementation on a major model to validate its approach for future funding rounds.
Source — Original announcement ↗
Worth a read?
Comments · 0