Transactions, signatures, API responses, and database invariants can be checked directly.
Verification infrastructure for agentic outcomes.
AI2Human compiles ambiguous requests into proof policies, evaluates judgment-based evidence through layered verification, escalates uncertainty, and emits auditable receipts that can safely trigger settlement.

A model score is not a verification protocol.
Once an agent workflow depends on screenshots, photographs, documents, identity-bound actions, or contextual claims, completion can no longer be inferred from plausible output alone. AI2Human treats this boundary as a protocol problem: specify what would count as proof, bind evidence to context, combine deterministic and probabilistic checks, expose uncertainty, and separate semantic judgment from economic settlement.
The open world has three epistemic zones.
Evidence requires task-conditioned interpretation, integrity signals, and explicit sufficiency rules.
When evidence cannot safely resolve the claim, a responsible verifier must abstain.
Generic LLM judges collapse all three into one prompt and one answer. The result is difficult to reproduce, calibrate, audit, or safely connect to payment.
Compile first. Verify second. Settle last.
A decision is a versioned function of policy and evidence.
Q = (intent, constraints, context, deadline, settlement_policy)P = C(Q) = (requirements, rules, capture, privacy, thresholds, version)E = (artifacts, claims, time, location?, identity?, integrity_commitment)O = D(P,E) ∪ F(E) ∪ M(P,E)d = Π(P,O,H) ∈ {pass, resubmit, manual_review, reject}R = Hash(Q,P,E_commitment,O,H?,d,versions,timestamp)The Proof Compiler is not a prose generator. It creates the ex-ante verification contract. Every required artifact maps to a rule; every decision traces to a versioned policy.
Selective automation is the safety mechanism.
Deterministic validation
Schema, required artifacts, deadlines, task binding, identity constraints, commitments, and uniqueness.
Evidence forensics
Hashes, container structure, EXIF/XMP, temporal confidence, editing signals, and batch consistency.
Semantic verification
Policy-specific multimodal judgments with structured reasons, confidence, and model provenance.
Selective escalation
Conflicts, integrity risk, provider degradation, and low confidence become human review—not forced verdicts.
The protocol assumes every participant can fail.
Blockchain settlement makes transfers auditable; it does not prove semantic truth. Multimodal models provide judgments; they do not become trusted oracles.
The receipt—not the campaign page—is the product.
- Policy
- proof-policy/v1 · immutable requirements
- Evidence
- SHA-256 commitment · scoped artifact references
- Observations
- deterministic + forensic + multimodal reasons
- Decision provenance
- models, thresholds, conflicts, escalation state
- Settlement
- eligible only after pass + idempotency validation
A third-party agent should be able to consume the verification result without entering AI2Human's website or trusting an unexplained private status flag.
Adjudicated cases become a controlled learning system.
Volume is not reliability.
False accepts, false rejects, calibration, and risk-coverage curves by task class.
Escalation, p50/p95 latency, model cost, reviewer minutes, and protected-value ratio.
Receipt completeness, audit reconstruction, reason agreement, and privacy burden.
The academic study compares a generic LLM judge with proof-policy, forensic, ensemble, and selective-verification variants. Independent raters and component ablations are mandatory before top-tier claims.
Implemented systems and research ambition are not the same thing.
Implemented
- Proof Compiler with deterministic fallback
- Structured evidence and integrity commitments
- Image metadata and container forensics
- Weighted multimodal review and escalation
- Payment idempotency and Base settlement
In development
- Stable standalone Verification API
- Privacy-scoped public receipts
- Reusable proof-policy registry
- Reviewer calibration metrics
- Expanded agent integrations
Research
- Task-specific reputation
- Memory-aware verifier routing
- Federated verification providers
- Portable receipt standards
- Domain-specific compliance policies
