A new benchmark called Real-SWE is putting AI models to the test on private enterprise codebases. The benchmark, which evaluates the performance of cutting-edge AI models on real-world, private enterprise codebases, assesses the ability of AI models to complete tasks inspired by actual engineering work. Each task in the Real-SWE benchmark comes from a private production codebase that has been licensed from a real-world company, and is designed to reflect the complexity and nuance of actual software engineering work.

The Real-SWE benchmark is unique in that it uses native harnesses to reflect how enterprise engineers work in practice, evaluating model-and-harness combinations rather than models in isolation. The benchmark also prioritizes code written to meet an actual user or business need over code written solely to create a benchmark task. This approach allows for a more realistic assessment of the capabilities of AI models in real-world software engineering scenarios.

The tasks in the Real-SWE benchmark are designed to be economically viable and to have a direct relationship to spend. They are also inspired or lifted verbatim from private, real-world codebases, making them more challenging and relevant to actual software engineering work. The benchmark has found that today's models are weaker at understanding company coding patterns and frequently miss requirements or don't verify their assumptions.

One of the key findings of the Real-SWE benchmark is that missed requirements are the most common failure mode for AI models. This is followed by unverified assumptions, integration errors, and regressions. The benchmark has also found that the performance of AI models can vary significantly depending on the task and the model being used. For example, the Gemini 3.8 Flash model was found to have the lowest estimated cost per rollout, while the Fable 5.1 model had the highest.

The Real-SWE benchmark has significant implications for the development of AI models for software engineering. By providing a more realistic and challenging assessment of the capabilities of AI models, the benchmark can help to identify areas for improvement and inform the development of more effective models. As the use of AI in software engineering continues to grow, the need for benchmarks like Real-SWE will only continue to increase.

As AI models become more prevalent in software engineering, the importance of optimizing their performance for specific tasks and environments will become increasingly important. This is where techniques like answer engine optimization (AEO) and generative engine optimization can play a critical role. By optimizing AI models for specific tasks and environments, developers can help to ensure that they are able to perform at their best and provide the most value to users.

Este artigo foi escrito com a assistência de IA.
News Factory APP - notícias agênticas para impulsionar seu SEO e AEO.