Skip to content
Intelligence RadarNews & Launches2 MIN READ

NVIDIA Announces SWE-Serve to Evaluate AI Coding Agents

NVIDIA has introduced SWE-Serve, a tool designed to evaluate AI coding agents on inference-engineering tasks, highlighting a significant gap between local tests and live serving.

Fathom Intelligence
Fathom IntelligenceFathom Layer Expert

NVIDIA has introduced SWE-Serve, a tool designed to evaluate AI coding agents on inference-engineering tasks, highlighting a significant gap between local tests and live serving. According to the announcement, SWE-Serve evaluates AI coding agents on 53 inference-engineering tasks derived from 83 merged SGLang pull requests across six families including model enablement, decoding, caching, scheduling, serving APIs, and distributed execution.

Nineteen tasks start a live server and test patches through the full serving path. Across these tasks, the same patches pass 45.9% of the time with complete verification but 69.4% when live-serving checks are excluded, indicating that about one in three patches that pass other tests fail live-serving validation.

Tasks spanning multiple runtime domains such as request handling, scheduling, model execution, and KV-cache management show a 21.3 percentage-point lower pass rate than single-domain tasks, with every tested model exhibiting the same gap.

Developed with input from the SGLang team, SWE-Serve evaluates the gap between passing local checks and working through the full serving path. The tool turns 83 merged SGLang pull requests into 53 executable tasks across six inference-engineering families, including speculative and advanced decoding, model and backend enablement, kernels, quantization, and performance, serving APIs and runtime correctness, caching and runtime state, and distributed execution and scheduling.

Twelve tasks run on CPU, while 41 use a single NVIDIA H100. This first release doesn’t evaluate other inference engines, multi-GPU execution, or multi-node serving.

Explore the SWE-Serve leaderboard and run SWE-Serve on GitHub to evaluate your coding agent on SGLang inference-engineering tasks.


Source: nvidia-developer

END OF SIGNAL

Don't Miss the Next Signal

Get high-impact hardware and AI launches distilled into your inbox. No noise, just the changes that matter.