Trust Ratings
Graded, comparable, and explained
A rating that cannot be interrogated is a marketing device. Every Deep Heuristics rating discloses its dimensions, its thresholds, and the evidence beneath each grade.
Dimensions
Seven dimensions, rated separately
We deliberately do not publish a single composite score. Collapsing these into one number conceals precisely the trade-offs a reader needs to see.
- Reliability
- Accuracy and calibration against the operating distribution, including whether stated confidence corresponds to observed correctness.
- Robustness
- Behaviour under distribution shift, edge cases, degraded inputs, and deliberate perturbation.
- Fairness
- Disparate performance and impact across relevant groups, assessed against the legal and ethical framing that applies to the use case.
- Explainability
- Whether a decision can be accounted for to the person it affects, and traced by those responsible for it.
- Security
- Resistance to adversarial manipulation, data exfiltration, and supply chain compromise.
- Privacy
- Data minimisation, memorisation and inversion exposure, and lawful basis for the processing performed.
- Governance maturity
- Whether the controls credited exist in practice: named ownership, working approval gates, and evidence that escalation has occurred when it should.
Interpretation
A rating is a statement about evidence, not about a vendor
Ratings attach to a system at a version, in a use context — never to a company. Two products from the same organisation can and often do rate differently, and a strong rating on one says nothing about the other.
- System-scoped, not vendor-scoped
- Version-bound
- Use-context specific
- Published thresholds
- Evidence-linked grades
- Re-rated on material change
Related
Deploy AI with confidence.
Start with an independent assessment of your highest-stakes AI system.