Global Decompression Benchmark
Compare decompression models against evidence, not reputation.
The benchmark evaluates documented models, tables, procedures and configurations against standardized historical exposure and outcome data. It is designed to show where evidence is strong, where performance changes with configuration and where the available data are not sufficient to decide.
What the benchmark evaluates
The benchmark separates model family, version, implementation and configuration. A result is attached to the exact system that produced it, rather than to a broad model name that may hide material differences.
Evaluation is constrained by the evidence available for each exposure. OpenDeco distinguishes calibration data from independent validation data and does not treat a counterfactual schedule as an observed outcome.
Explore the current release
No public release yet.
Views in the explorer
- 1Overview
- 2Models
- 3Datasets
- 4Results
- 5Failure regions
- 6Sensitivity
- 7Methods
- 8Downloads
- 9Changelog
How to read a benchmark result
A benchmark score is not a universal safety ranking. Performance depends on the dataset, outcome definition, model configuration, calibration history and question being asked. OpenDeco therefore publishes the result together with its scope, uncertainty and failure regions.
Identity
Check the exact model version, implementation and configuration before comparing results.
Evidence
Check whether the evaluated dataset helped create or calibrate the model. Training-data performance is not independent validation.
Failure regions
Average performance can conceal exposure regions where a model behaves differently. OpenDeco keeps those regions inspectable.
Reproduce the release
Every public benchmark release carries the dataset versions, model identities, code commit, run manifest, protocol version and artifact checksums needed to understand what was executed. Where redistribution or proprietary restrictions prevent exact reproduction, the limitation is stated explicitly.