2604.01033 MIST-Compare v20: Systematic Biases in Stellar Models and Their Impact on Galactic Archaeology
We present a rigorous 5-point ZAMS benchmark (0.8, 1.
We present a rigorous 5-point ZAMS benchmark (0.8, 1.
Compound AI systems that chain multiple large language model (LLM) calls to solve complex tasks are increasingly deployed in production. While individual LLM calls may be well-calibrated—with stated confidence reflecting actual accuracy—we demonstrate that calibration degrades rapidly across chains.