NEW The 2026 State of AI Trust Report is now available. Read the report →

Security Research

Understanding the attack surface as it forms

Original research on adversarial technique, agentic system risk, and the supply chain beneath modern AI — conducted to sharpen our assessments and published to raise the floor.

Focus areas

Where we concentrate

Injection and instruction integrity Direct, indirect, and multimodal prompt injection, and the architectural patterns that make instruction-data separation tractable.
Agentic risk What changes when a model can act: tool boundaries, memory poisoning, multi-agent trust, and blast radius.
Supply chain Provenance and integrity of weights, adapters, datasets, and the packaging and registries around them.
Evaluation integrity Benchmark contamination, evaluation gaming, and why a strong published score is weak evidence.

Taxonomies we work from

Shared vocabulary, so findings are comparable

OWASP Top 10 for LLM Applications
The 2025 edition: prompt injection, sensitive information disclosure, supply chain vulnerabilities, data and model poisoning, improper output handling, excessive agency, system prompt leakage, vector and embedding weaknesses, misinformation, and unbounded consumption.
MITRE ATLAS
Adversarial tactics and techniques observed against AI-enabled systems, used to structure red team objectives.
NIST AI RMF
The Measure and Manage functions supply the structure for converting a security finding into a governed risk decision.
CVE and CWE
Conventional vulnerability identifiers still apply to the substantial non-model surface: serving infrastructure, orchestration, and integrations.
Data protection and cybersecurity controls represented as layered digital safeguards.

Disclosure

We publish after the fix, not before

Research that identifies a weakness in an identifiable product goes to the vendor first, with a defined remediation window before publication. We publish regardless of whether a fix arrives — but never before the window closes.

  • Vendor notified first
  • Defined remediation window
  • Coordinated publication
  • No exploit code for live systems
  • Credit to reporters
  • Publication regardless of outcome
Read the disclosure policy

Deploy AI with confidence.

Start with an independent assessment of your highest-stakes AI system.