Augment, a startup focused on AI‑assisted software development, announced that its context engine delivers the same accuracy as Anthropic’s Claude Code while using 33% fewer tokens, according to a benchmark the company ran on the Terminal‑Bench suite. The result, the firm says, stems from a retrieval system tuned specifically for large code repositories.

“We ran Terminal‑Bench with Claude Code and Augment Code, the same model, and we completed at similar accuracy, but we were 33 percent more efficient,” said Augment’s co‑founder during a recent interview. The efficiency gain, he explained, comes from asking the model the right question and feeding it precisely the code snippets needed, rather than letting it wander through irrelevant context.

Anthropic, the creator of Claude Code, has previously reported that adding semantic code‑navigation tools does not produce measurable evaluation gains. The discrepancy, Augment’s representative suggested, is rooted in the differing retrieval engines each organization uses. “Not all retrieval systems are equal, just like not all databases are equal,” he said, emphasizing that the quality of the context engine matters as much as the language model itself.

Augment’s team spent roughly 18 months, beginning in 2022 before the ChatGPT boom, researching retrieval and embedding models tailored for massive codebases. Their effort focused on determining which pieces of code should be placed in the embedding space to maximize relevance and speed. The company claims this research is baked into its current retrieval models, enabling rapid identification of the most useful code fragments.

When asked why some users see no benefit from Claude Code with a retrieval‑augmented generation (RAG) setup, the Augment executive replied that implementation details matter. “If someone tried Claude Code with a RAG implementation and didn’t see benefits, that’s because the implementation and the context engine are very, very different,” he said.

Industry observers note that the debate underscores a broader challenge: evaluating AI coding assistants when the surrounding infrastructure varies widely. While Anthropic’s internal metrics may not capture the advantage of a custom retrieval layer, Augment’s benchmark suggests that a well‑engineered context engine can meaningfully reduce token consumption—a key cost factor for developers using paid API calls.

The claim, if verified across more real‑world projects, could influence how companies integrate AI into their development pipelines, prompting a shift toward building or licensing specialized retrieval components rather than relying solely on the language model’s native capabilities.

Este artigo foi escrito com a assistência de IA.
News Factory APP - notícias agênticas para impulsionar seu SEO e AEO.