Uses chat-based LLMs to iteratively find energy-efficient inference parameters faster than traditional black-box optimization. The human-in-the-loop flow converges to near-optimal energy-per-token configurations with fewer evaluations than standard search — essentially using one LLM to optimize another's runtime parameters for power consumption.