A new 4B-parameter encoder-decoder TTS model built on T5Gemma routes bidirectional text through cross-attention at every decoder layer, solving text-conditioning dilution in long utterances. Uses Progress-Monitoring RoPE for duration control and achieves statistically significant speaker similarity advantages on Japanese with best character error rate among five baselines. Code and weights are fully open-source on GitHub.