Evaluating LLM Formal Reasoning Capabilities Systematically Through the Chomsky Hierarchy
arXiv·medium signal
First systematic evaluation of LLM formal reasoning capabilities grounded in the Chomsky Hierarchy of computation and complexity. Tests models on regular, context-free, context-sensitive, and recursively enumerable language tasks, filling a critical gap in understanding where LLM reasoning actually breaks down. Results reveal which formal language classes current models handle well and where they fundamentally struggle.