Sitemap

Even the Top LLM Failed to Reliably Flag Some Risks Related to 40% of Safety Facts

3 min readAug 29, 2025

In a recent study, even the best-performing LLM failed to reliably flag safety risks related to about 40% of common safety facts tested. CDS master’s student Yueh-Han (John) Chen and his collaborators CDS PhD student Guy Davidson and CDS Associate Professor Brenden Lake discovered this troubling gap when they tested leading AI systems on their ability to generalize well-established safety facts to new situations.

Their new benchmark, detailed in the paper “SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts,” revealed that Claude 3.7-sonnet, the top-scoring model, reliably identified safety concerns for only 58% of safety facts. Other leading models performed far worse — OpenAI’s o1 managed just 41%, while Meta’s Llama-3–70B passed a mere 6% of safety tests.

“Most current systematic generalization benchmarks are synthetic and don’t reflect real-world scenarios,” Chen explained. “We needed a benchmark that measures generalization abilities while demonstrating real-world consequences.”

The researchers designed scenarios where users might unknowingly request dangerous advice. In one example, a parent asks for creative snack ideas while mentioning they’re considering slicing sausages width-wise for their 18-month-old — a choking hazard according to CDC guidelines. While some models provided helpful food suggestions, they failed to warn about the specific danger mentioned in the prompt.

The team built their benchmark around 104 safety facts sourced from reputable organizations including the CDC, FDA, and American Academy of Pediatrics. These facts span seven domains: child safety, animal care, chemical handling, outdoor activities, medicine, senior care, and cybersecurity. Each fact was then systematically augmented to create over 10,000 test scenarios, all validated by human annotators. A fact was considered “reliable” only if the LLM could identify the hazard in all (or nearly all) of the many scenarios it appeared in.

The research revealed a particularly concerning finding: model capability and safety performance showed weak correlation. “We use Chatbot Arena — which aggregates human judgment — as a proxy to measure model capabilities,” Chen said. “But we found the correlation between capability and safety to be just 0.2, which is very weak.”

This disconnect suggests that simply building more powerful AI systems won’t automatically make them safer. Models that excel at complex reasoning tasks can still miss obvious safety warnings when those warnings require applying known facts to novel situations.

The benchmark also exposed how context affects safety awareness. Models performed worse with longer prompts containing hidden safety concerns, achieving only 80% responses that are safe with 100-word contexts compared to 90% with simple instructions. User tone mattered too — depressed emotional contexts reduced safety performance to 86%.

Chen began this research project before even starting his master’s program, reaching out to professors after being accepted to CDS. “I went through the faculty contact list on the CDS website and found Brenden’s research really compelling,” he recalled. His initiative paid off, giving him an early start in AI safety research.

The work represents the first benchmark to test systematic generalization of safety knowledge in real-world scenarios. Unlike previous safety evaluations that focus on explicitly harmful requests, SAGE-Eval targets the subtle but dangerous failures that could emerge when naive users interact with AI systems.

The researchers recommend that AI companies incorporate SAGE-Eval into pre-deployment testing. They also developed forecasting methods to predict how safety performance might degrade as models encounter even more diverse user prompts in real-world deployment.

“Frontier LMs still lack robust generalization ability,” the researchers concluded. Their findings suggest that achieving truly safe AI will require more than just scaling up models — it will demand new approaches specifically designed to ensure that AI systems can reliably apply safety knowledge across all the varied ways humans might interact with them.

By Stephen Thomas

--

--

NYU Center for Data Science
NYU Center for Data Science

Written by NYU Center for Data Science

Official account of the Center for Data Science at NYU, home of the Undergraduate, Master’s, and Ph.D. programs in Data Science.