BLUEX v2 creates a public dataset from 2022-2025 UNICAMP and USP second-phase exams and evaluates 21 LLMs via LLM-as-judge, finding scores from 4.18 to 9.10 with math and image understanding as weakest areas.
Sabi\’a-4 technical report
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CL 3years
2026 3representative citing papers
Creates the first bilingual clinical benchmark from Brazilian cases and reports that English performance advantage exists only in diagnosis retrieval, disappearing in the other three tasks.
MARCA is a bilingual benchmark using 52 questions and validated checklists to evaluate LLM web-search completeness and correctness in English and Portuguese.
citing papers explorer
-
BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams
BLUEX v2 creates a public dataset from 2022-2025 UNICAMP and USP second-phase exams and evaluates 21 LLMs via LLM-as-judge, finding scores from 4.18 to 9.10 with math and image understanding as weakest areas.
-
Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese
Creates the first bilingual clinical benchmark from Brazilian cases and reports that English performance advantage exists only in diagnosis retrieval, disappearing in the other three tasks.
-
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
MARCA is a bilingual benchmark using 52 questions and validated checklists to evaluate LLM web-search completeness and correctness in English and Portuguese.