Meta Makes Its Mark in AI Coding with a Surprisingly Robust Agent

Meta has made its entry into the realm of AI coding tools, joining the ranks of established players like Anthropic’s Claude Code and OpenAI’s Codex. The company has introduced a beta version of Muse Code, which aims to deliver competitive performance at an accessible price point.
This terminal-based tool is powered by the Muse Spark 1.2 model and is designed to manage “complete” tasks, even within extensive codebases, as outlined by Meta CEO Mark Zuckerberg. Muse Code is capable of planning modifications, generating code, and verifying outcomes. It utilizes background agents to maintain context, while larger projects are delegated to separate “sub-agents” that operate independently.
Every interaction with the model, including tool usage, is logged, ensuring that the agent can resume its work seamlessly in case of a failure.
One of the appealing aspects of Muse Code is its low cost, as mentioned by AI lead Alexandr Wang in an interview. Users have the option of a pay-as-you-go model, charging $1.25 for every million input tokens and $4.25 for a million output tokens. Alternatively, a contributor tier is available, which Wang describes as being “more than 10 times cheaper.” This option is geared towards those willing to assist in enhancing the model, making it particularly well-suited for newcomers or team settings.
Meta is also accommodating requests for data privacy, providing options for organizations that wish to restrict their sensitive information from being utilized in AI training processes.
How does Muse Code stack up against Claude Code and Codex?
Best performance within Meta’s ecosystem
Muse Code serves as a coding framework that allows for the management of several models now but is optimized for use within Meta’s platforms. Wang indicated that Muse Spark 1.2 will achieve optimal results, as it was specifically developed alongside the agent.
Under these conditions, Meta asserts that Muse Spark 1.2 competes effectively with other models. It surpasses GPT 5.6 Terra in the Terminal-Bench 2.1 software engineering benchmark, while also performing closely to Claude Opus 5. In the long-horizon DeepSWE 1.1 test, Muse Code ranks behind both but still performs better than X.ai’s Grok and Google’s Gemini 3.6 Flash.
Using Meta’s in-house coding assessments, the model positions itself between Claude Opus 5 and GPT 5.6 Terra.
While speed may not be the primary factor for selecting Muse Code at this point, it is noteworthy that Meta demonstrates competitive capabilities in coding within just four months of launching the initial Muse Spark model. Their pricing model may also attract users in specific situations, although regular use of AI coding tools could lead to higher subscription costs.



