Kog Dives Deeper to Maximize GPU Inference Efficiency

The competition to enhance AI inference speeds is heating up, highlighted by Cerebras’s successful IPO in May. Meanwhile, the French startup Kog is exploring the untapped capabilities of conventional GPUs.
In May, Kog attracted attention on Hacker News by demonstrating that “exceptionally fast single-request decoding can be achieved using standard data center GPUs that companies typically have,” specifically showcasing the AMD MI300X and Nvidia H200 in their test.
While some were disheartened to learn the technology does not apply to typical laptop GPUs, others were intrigued. With AI inference speed and cost being significant constraints, Kog’s ambition to enhance existing hardware through software optimization captured considerable interest. CEO Gaël Delalleau shared that they have already generated “200 meaningful business leads.”
Initial responses suggest that software engineering will be the first area of application. Regular users of Claude Code are familiar with prolonged wait times for results. Notably, Anthropic acknowledges the value of speed, offering a premium for Claude’s Fast Mode.
Kog aims to attract customers frustrated by these delays, typically those who depend on AI for professional tasks. The startup is also collaborating with design partners, enabling users to create games and applications via prompts. Delalleau emphasized that faster results from the Kog Inference Engine (KIE) could directly increase revenues for these partners.
The company recognizes that the market is still evolving. Through its observations, Kog discovered that potential customers are often reluctant to fine-tune smaller models. “Thus, we have dedicated ourselves to accelerating the development of larger models in response to observed demand since our launch,” Delalleau noted.
This commitment presents a significant challenge for Kog as it strives to fulfill its promise of delivering “30x quicker LLM inference.” The demonstration highlighted an impressive performance of 3,000 tokens per second (TPS) for each request, utilizing a specialized smaller model—the now open-sourced Laneformer 2B, which contains around 2 billion parameters.
Countering doubts from critics, Delalleau is optimistic that this methodology will be just as effective with larger language models, which often pose difficulties for inference chips due to their size. “The future looks bright for GPUs,” he asserted, explaining that the belief that they aren’t ideal for decoding is a misconception, given the increasing memory bandwidth of newer models waiting to be tapped into.
Kog shares the sentiment with ZML, another French entity, which has developed hardware-agnostic software that circumvents Nvidia’s CUDA for quick inference across various chips. However, Delalleau claims that Kog aligns more closely with Hazy Research from Stanford, focusing on deeper GPU acceleration.
Though Delalleau is not a researcher, and his previous startup, Stribe, founded in 2009, differs from Kog, he collaborates with former co-founder Kamel Zeroual, whose firm, Varsity VC, recently co-led Kog’s seed funding. Kog’s unique approach is influenced by Delalleau’s background.
Having studied solid-state physics at École Polytechnique in France, he transitioned into offensive cybersecurity, also known as white-hat hacking. Delalleau believes this experience fosters a mindset that encourages his team to learn the principles of physics and GPU functionality for maximum efficiency.
In terms of hacking, the multiple-time DEFCON CTF tournament finalist explained that it taught him “to reverse-engineer at a grassroots level—understanding assembly language and binary code to grasp functionality and use it creatively.”
However, this hands-on approach can be time-consuming. “We dedicate several weeks or even months for each new GPU to thoroughly investigate and perform engineering research on that hardware.” With a compact team of 11, this limits the number of chips Kog can engage with in the near term.
Long-term, Kog envisions evolving its methodologies into agent-based pipelines capable of supporting numerous chips and models. As Europe seeks to enhance its capabilities in these areas, this trend could provide advantageous support for Kog, which is backed by Scaleway and funded through initiatives from France’s Bpifrance and French Tech 2030 program.
At this stage, Kog needs to demonstrate its approach’s effectiveness with LLMs, a crucial step for securing additional funding. “Once we implement our first significant model with a 10x speed increase—which I anticipate will happen in September—we can start showcasing customer traction and consequently aim for our Series A funding,” Delalleau stated.



