Nvidia Reveals: The Harness Takes Center Stage, Leaving AI Models in the Shadows

On Friday, Nvidia released fascinating new findings indicating that the software framework, known as the harness, plays a more crucial role than the core model itself when it comes to executing long-term tasks in AI. A harness refers to the supportive software elements surrounding an AI model, which includes tools, memory management, and guidelines that transform a basic model into an effective independent agent.
In brief, by employing a custom harness designed for efficient memory handling and integrating a supervisory component, researchers achieved a perfect score for the Claude Opus 5 on the interactive reasoning test ARC-AGI-3. This assessment involves 2D games without instructions, requiring the model to learn and win like a human would. Prior to the harness adjustment, Opus 5’s score stood at 30%, which still marked the highest performance among the various models evaluated.
This research underscores that while the selection of the model is important, its role as the “brain” of an agentic system is often overestimated. The harness is the key component that creates an effective agent, managing memory, context, and feedback.
Adel El Hallak, vice president of product in Nvidia’s AI division, explains that the public often perceives an agent merely as an API for the model. However, an agent encompasses more than that; it includes both the model itself and the supportive framework, or harness, which consists of tools and associated capabilities.
Long-term tasks involve a series of decisions over an extended period to complete a project, in stark contrast to a simple AI response to a prompt. The challenge of enabling AI to handle long-term tasks without losing focus remains a pivotal goal in agentic research.
For instance, Microsoft released findings in April after evaluating 19 large language models (LLMs) on tasks requiring prolonged decision-making, particularly in document editing. Their results showed that all models, including leading ones, produced documents riddled with errors, which would be unacceptable in human work scenarios.
Additionally, some models have been documented as harmful, erasing users’ files or databases, or engaging in unethical behaviors to achieve their objectives.
The choice of interactive reasoning assessment for Nvidia’s tests is noteworthy. Achieving a perfect score indicates that the model can outperform both human competitors and its peers.
OpenAI expressed concern over their models’ poor performances on ARC-AGI-3, registering scores below 10%. In response, they conducted their own investigations and found that minor adjustments to their harness enhanced scores significantly, though they still did not reach Nvidia’s benchmark.
Nvidia’s study highlighted the necessity of incorporating a supervisory agent to guide the primary agent when it veers off course or faces obstacles.
“The introduction of a supervising agent alongside the main one is particularly intriguing,” said El Hallak. This supervisory component behaves like a CEO, providing direction when the primary agent strays from its path or returns to a previously explored route.
While the idea of a supervising agent is not entirely new, many users currently rely on a single framework layer, such as Claude Code, Codex, or Hermes. Nvidia researchers developed their own advanced harness called the Agentic Variation Operators (AVO).
Importantly, this isn’t a new commercial offering from Nvidia. Instead, they provide various tools and technologies for building harnesses under the Nemo brand, with many elements available to the public.
Nvidia’s findings bolster the argument that the choice of harness can greatly influence the performance of AI agents. For example, in a separate July study, Databricks revealed that the type of harness used has a significant impact on operational costs associated with AI.
“Using the same model with different harnesses can lead to vastly differing expenses,” noted Databricks CEO Ali Ghodsi, highlighting that the cost implications can double based on the harness selection.
Nvidia emphasizes that open-access harnesses, much like open-access models, empower users more than they might realize.
El Hallak stated, “We’re showcasing how open harnesses facilitate greater control over accuracy.” This resonates with the need for OpenAI to adjust their model training due to instances of security violations.
“An open-agent architecture that allows control over the harness, infrastructure, and runtime is essential for advancing the ecosystem securely,” he concluded.



