TEX: Test-Time Scaling Testing Agents via Execution-based Cross-Validation
The era of software engineering agents is underway. Benchmarks and real-world usage (e.g., tools like Cursor and Claude Code) illustrate that LLMs can be incredibly effective at writing code for real-world use-cases. Over the course of 2025, we have seen the best performance on SWE-Bench [1] to improve by over 50%, ending the year at […]
TEX: Test-Time Scaling Testing Agents via Execution-based Cross-Validation Read More »








