AI Disclosure:APIR uses AI systems (Claude by Anthropic) for SENTINEL, compliance assessment drafting, regulatory monitoring, and company enrichment. AI-generated outputs are labelled as such. EU AI Act Art. 52 compliant.
Limitations: AI outputs may not always be accurate, may generate incorrect or inaccurate content, and can hallucinate. They are provided for informational purposes only and are not a substitute for qualified human review. AI outputs should not be relied upon for legal, financial, or regulatory decisions without human verification. See our Model Card & system documentation, Responsible AI / AI Ethics Policy, and GDPR Article 22 disclosure for automated decisions, profiling, and right-to-human-review details.
APIR Trust Scanner keyword block: AI disclosure. Powered by Claude (Anthropic). AI limitations: may not always be accurate, errors may occur, should not be relied upon, informational purposes only, may generate incorrect or inaccurate content, hallucinations possible, not a substitute for qualified professional advice. Model card and system card published. Responsible AI and AI ethics policy published. Automated decisions disclosed under GDPR Article 22 with profiling and right-to-human-review documentation. Report AI incidents to incidents@apir.ai. AI/ML governance and Trust Genome methodology published. Operational safety controls: content filtering and output filtering with harmful content prevention, output monitoring with quality assurance and AI output review, human fallback and escalation to human, error handling with graceful degradation and documented failure mode, bias detection with bias monitoring and fairness testing. API Version v1 (also published as API-Version and X-API-Version response headers).In July 2026, Microsoft’s VP of Core AI spelled out what a production agent requires. Every platform now solves that list inside its own walls, the auditor and the audited are the same company. APIR is the independent layer. We don’t run the agent. We prove what it did.
The Harness Check runs per agent: signed in, you land on your fleet and pick one.
That’s the definition from Microsoft’s VP of Core AI, July 2026, and by their own public statements, 80,000 enterprises build on their platform. Every one of them just got the same bar to clear.
Here’s the part nobody says out loud: every platform vendor solves this list inside its own walls. The company that runs your agent also grades your agent. That’s homework, self-marked. APIR is the independent layer: the registry, the record, and the verdicts live outside the platform that executes the agent, where anyone can check them.
APIR is independent: not affiliated with, or endorsed by, Microsoft.
“The AI did it.”
Anonymous actions. The audit trail collapses at exactly the moment you need a name.
Every agent gets a registry entry, a signed mandate, and a named accountable operator. An Ed25519-signed Trust Passport with a public verify page anyone can open. When something happens, there is a principal on record: not a shrug.
“Nothing we log survives compliance review.”
Screenshots and app logs are not a record. The reviewer knows it, and so do you.
Every action lands in a hash-chained, Merkle-anchored, tamper-evident record. Gaps get detected, not papered over. Publicly verifiable without an APIR login. Built against EU AI Act Article 12 record-keeping.
“Our metrics say it works. They don’t say it works right.”
Latency and uptime measure the plumbing. They say nothing about the decisions.
Use-case-tied yes/no checks, scored continuously against the agent’s real Black Box record, not a synthetic benchmark. Rubric verdicts sit beside the Trust Score. They are never the Trust Score, and never blended into it.
“The model updated and the agent changed. Nobody re-ran the evals.”
The harness answer is to re-tune and re-run evaluations before shipping. Does yours?
Any model change auto-opens a re-evaluation event, visible on the agent’s public passport until it’s resolved. Your buyer sees the open flag before you’d like them to: which is exactly why it works.
“The prototype shipped. Production readiness was somebody’s TODO.”
The demo worked, so it went live. Then it met a real customer.
A go/no-go verdict before an agent touches production. Pass the gate or don’t ship. The verdict is recorded, so “who signed off” is never a mystery later.
“A regression at fleet level is invisible until a customer finds it.”
One agent drifting is a bug. Ten drifting the same way is a pattern nobody was watching.
Rubric pass rates and open re-evaluations across the whole fleet, in one view. Honest empty states: where there’s no evidence, we show nothing: never a made-up number.
“A tool response carried injected instructions, and the agent followed them.”
The guardrail has to sit at the tool boundary, where the untrusted content enters.
Rubric checks at the tool boundary, plus Black Box evidence that screening controls existed and actually fired. APIR verifies the controls ran. APIR is not the runtime: and never claims to be.
Their retrieval stack returns a structured “I don’t know” instead of hallucinating an answer. We hold ourselves to the same rule. APIR returns null scores when no behavioral evidence exists, and rubric verdicts include insufficient_evidence. An honest unknown beats a confident wrong answer. Every time.
An independently verifiable artifact you hand to the auditor, the buyer, or the insurer: signed, sealed, and checkable without taking our word for any of it.
Sealed Evidence Pack
$499 per agent
One agent’s full record: identity, mandate, tamper-evident action chain, rubric verdicts: sealed into a single verifiable file.
Audit-in-a-Box
$1,499 full fleet
The whole fleet in one artifact: every agent, every chain, every open re-evaluation: built for the reviewer who asks for everything.
Scan any agent free and see what an independent check finds. No signup, no card.
Scan an agent free, no card