Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark
CAST is a benchmark for evaluating context-conditioned word-level stress in text-to-speech. We show that while language models reliably infer appropriate emphasis from discourse context, TTS systems frequently fail to realize it in speech. Alongside the benchmark, we release an extended generation pipeline to support future research on context-aware speech synthesis.