Research
Five Frontier Models Benchmarked Against GraphWalker for Model-Based Test Path Generation
This empirical evaluation pits GPT-5.1, GPT-5.2, Claude Opus 4.5, Claude Sonnet 4.5 and Gemini 2.5 Pro against the state-of-the-art model-based testing tool GraphWalker and its built-in random and quick-random algorithms for edge and vertex coverage. It uses four GraphWalker models of escalating complexity, two web applications (Parabank, Testinium) and two hardware ones (TLC, RISC-V). The reported result is strong potential for LLMs to optimize and shorten both test paths and step sizes, which matters because path length is the scalability barrier that has kept MBT out of industrial adoption.
↳ Follow the thread