About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language
Project Overview Help shape the future of AI in Brazil! We are building a next‑generation benchmark to evaluate how well AI systems handle complex, real‑world questions in Brazilian Portuguese . This project will help improve how