Selection Scope
Only completed simulations with available analyses can be selected for cross-simulation benchmarking.
Compare semantic network metrics across completed crisis simulations.
Ready for Benchmarking
Select two or more simulations from the sidebar to calculate aggregated metrics across semantic networks.
Common criteria for selecting completed audits, aggregating model metrics, and comparing semantic network behavior.
Only completed simulations with available analyses can be selected for cross-simulation benchmarking.
Metrics are grouped by AI model across the selected simulations and summarized with mean and standard deviation.
Semantic network metrics, CASA adherence, null-model baselines, and assortativity indicators are compared across selected audits.
The exported CSV preserves model-level means and standard deviations for every benchmark metric shown in the matrix.
Interpretation: Use this view to compare model behavior across simulations, not to inspect a single audit in depth.