AI-Enhanced Decision Making Failures
AI systems advising or making high-stakes decisions in healthcare, criminal justice, and infrastructure may encode biases, fail catastrophically on distribution shift, or be misused by bad actors.
Scale of the Problem
AI medical diagnostic tools show 20-40% error rate disparities across racial and socioeconomic groups.
Cross-cutting. Direct harms include medical misdiagnoses, flawed parole and sentencing decisions, erroneous military targeting, and flash-crash-style financial failures. Indirect effects on political and strategic decisions influence every other catastrophic risk. Small biases at scale can produce large aggregate effects over many decisions.
Why This Is Pressing
Organizations are rapidly deploying AI in high-stakes decisions across medicine, finance, criminal justice, and military targeting. Failures can happen silently and compound at scale. AI advice increasingly shapes decisions about AI itself - what to build, how to regulate, how to deploy - creating feedback loops where errors in decision support become errors in the overall trajectory of the technology. Improving the calibration, auditability, and integration of AI in human decisions is a tractable lever on many other risks.
Why It's Neglected
AI evaluations and auditing is a small field, with roughly 50 million dollars per year across organizations like METR, Apollo, and academic labs. Most lab alignment work focuses on worst-case misuse rather than everyday reliability. Public-sector evaluation capacity remains limited.
Can It Be Solved?
Dangerous-capability evaluations, red-teaming, uncertainty calibration, forecasting tournaments, and decision-support tooling are all showing progress. Standards such as the NIST AI Risk Management Framework provide scaffolding. Forethought, ARIA, and the UK and US AI Safety Institutes are funding related work.
Research & Solutions
Algorithmic Accountability: When AI Gets High-Stakes Decisions Wrong
WorldProblems Solved
Sign in to read full documents, vote on solutions, and submit your own research.
What You Can Do
What this problem still needs
Research at METR, Apollo Research, Forethought, or Epoch. Policy work on AI evaluations, auditing, and procurement standards. Forecasting roles at Good Judgment or Metaculus. Build AI tools that improve decision quality in medicine, law, or policy. Work on calibration, uncertainty quantification, and auditability.
Priority Score
20
Breakdown
Fund this problem
Support the researchers and organizations working on AI-Enhanced Decision Making Failures.
Donate to this causeContribute to solving this
Create a free account to submit research, vote on solutions, and follow this problem.
Join for free