Revolutionary AI Model Unveiled by ChatGPT Creator Sparks Excitement Among Developers

Diogo Almeida faces disappointment despite groundbreaking achievements in AI.
Diogo Almeida, a former researcher at OpenAI, played a significant role in developing the chatbot and creating reinforcement learning from human feedback (RLHF), a pivotal training approach that has shaped artificial intelligence today. Despite these advancements, Almeida expresses dissatisfaction with the outcomes.
“We have something amazing, yet it lacks practicality,” he stated in an interview. “I’ve struggled with this realization: we’re focusing too heavily on human language… While our proficiency in processing language has improved over the years, it’s not effective for automation since machines operate on a different level.”
Two years back, Almeida departed from OpenAI to launch TypeSafe AI, a startup dedicated to addressing this challenge. Recently, they unveiled a new transformer-based model called Jev, which diverges from large language models (LLMs). Instead of generating text, Jev provides probabilities, referred to as “calibrated decisions.”
This approach that avoids conventional language has multiple benefits: it offers rapid and cost-effective processing while eliminating the potential for hallucination by allowing users to preset outcomes. Furthermore, its output tokens are free, and input tokens are measured in billions, making it an affordable option.
There’s significant interest from developers; in fact, demand for Jev was so high that the company temporarily struggled to keep up with API requests. Many developers find Jev particularly advantageous for software automation, perceiving it as a cost-effective and robust means of enhancing their projects.
For instance, Pranit Sharma from Vercel reported that the company previously used OpenAI’s ChatGPT Luna for command safety classification. After switching to Jev, they observed a performance increase of 5 to 18 times in both speed and accuracy.
Another user, Bryo AI’s Nikhil Mudholkar, compared Jev with Gemini for categorizing business emails. While Gemini proved to be marginally more accurate, it was also 10 to 20 times more expensive. Mudholkar appreciated Jev’s confident scoring, noting, “it uniquely offers a genuine probability, making it ideal for automating tasks!”
In addition to replacing LLMs for certain functions, Jev can also enhance them by acting as a safeguard against errors. Using Jev for monitoring other agents can help manage costs, with Almeida suggesting it as a sensible solution. He envisions its deployment for tracking LLM agent activities and preventing misuse.
“Essentially, it shifts the hallucination responsibility partially to the user,” observed Armin Ronacher, CTO of Earendil, which develops the open-source model harness Pi. “The user must interpret outcomes based on confidence levels—deciding whether to act on a 50% probability or trust a 95% likelihood.”
Ronacher also identified model routing as another potential use for Jev. Determining whether a specific workload needs a particular model would be beneficial, and Jev’s efficiency could make this real-time sorting feasible.
Almeida’s ambition is to make intelligence more accessible, inspired by the economic theories of William Stanley Jevons, whose paradox suggests that reduced costs lead to increased usage. Almeida believes that decreased intelligence costs will result in broader adoption across various applications.
“We envision intelligent software permeating everyday life, evolving in a decentralized manner, much like the early days of the internet, rather than through the mega applications currently in vogue,” Almeida explained.
While details about Jev’s architecture remain undisclosed, some industry watchers speculate it builds on existing open-weight LLM frameworks. The company characterizes Jev as a “System One model,” emphasizing intuition over extensive reasoning, and is specifically designed for particular tasks. Almeida mentioned that Jev has been developed entirely using synthetic data through a method he dubs “reinforcement learning from calibrated decisions.”
“We confidently bet on generating our own data, which has proven to be one of the wisest decisions I’ve ever made—more valuable than our launch and RLHF,” he stated. “Half of our operations focus on mastering this domain of well-established synthetic data, which has now become a passion of mine.”
Currently, Jev stands as a unique offering, but Ronacher anticipates that competitors may emerge now that its value is recognized.
“In hindsight, we should have identified this earlier. However, LLMs are often economical and heavily supported, which lessens the need for creativity,” he commented.
TypeSafe intends to explore more versions of Jev across additional modalities. When asked if TypeSafe functions as a cutting-edge research lab, Almeida remarked, “the main product of conventional labs can often be fear or hype. My goal is to produce genuine intelligence…though we won’t be a lab that bets on limitless wealth or aims to develop a deity in a data center.”



