Aggregate human scoring of AI-generated reports across data, systemic risk, and action planning.
Averages always show their review sample and coverage. Observed ranking is descriptive, not a benchmark.