RocSite AI Evaluator audits the clinical AI tools your hospital deploys, with cryptographic provenance and zero vendor incentive bias.
Available by request to qualified institutional evaluators. Not a consumer product, not an AI vendor, not for individual case-by-case clinician use.
Hospitals deploy clinical AI tools. The vendors who sell those tools validate their own products. Regulators set baseline approval criteria but cannot audit the model's behavior on the actual patient population it sees once it's installed. That leaves AI Governance Committees, the people inside the institution who are accountable for AI decisions, relying on manual case-by-case audit processes that don't scale.
These are real failure modes that exist in deployed systems today: a small subdural hematoma isodense to cortex that a narrow-vision detection model misses because its training cohort didn't emphasize that presentation. A hyperdense meningioma flagged as a hemorrhage because the texture pattern matches what the model was trained to call positive. Drift caused by a new CT kernel installed last quarter, or by a patient population gradually diverging from the cohort the model was trained on.
None of these are fixed by reading vendor documentation. They show up only when somebody independent looks at the discordant cases, the cases where the AI and the radiologist disagree, and explains why. That work is the gap. The AI Evaluator fills it.
The Evaluator reviews de-identified cases where AI predictions and human reads disagree, on the bench, off the patient record, never in the live worklist. It runs only on the discordant set, the cases where the audit signal is, and produces structured, clinically grounded explanations for each one as audit records for a governance file.
Identifies AI-vs-human disagreement and focuses analysis there. Concordant cases skip expensive review.
Surfaces likely failure modes grounded in named clinical criteria sets (Fleischner 2017, LI-RADS, BI-RADS, Lung-RADS, TI-RADS, Sepsis-3 / SOFA / qSOFA, CHA2DS2-VASc, NIHSS, Ottawa Ankle, ASIA). Clinical-sounding claims that aren't grounded in a named criteria set are flagged in the audit record, not buried in the prose.
Multiple imaging models run against the same scan. Divergence patterns surface what a single-model audit cannot.
Tracks deployed model performance over time against the vendor's stated performance characteristics (including 510(k) summary tables where published). Alerts when adjudicated discordance rate exceeds the institution's configured threshold, or, where ground-truth labels are available in a parallel adjudication workflow, when measured sensitivity or specificity drifts outside vendor-stated bounds.
A queryable database of identified failure patterns with remediation guidance. Vendors can use it to avoid known modes ahead of deployment.
Every evaluation gets an evidence hash, Ed25519 signature, timestamp, and methodology version stamp. The signature on a signed evaluation can be verified independently against our published public key.
RocSite does not sell clinical AI tools. We evaluate them. The structural independence is the asset. Vendors approach us regularly, asking for integration into their stacks or licensing arrangements. We decline, because the moment a validator sells anything to the systems it validates, the validation loses meaning.
We may publish anonymized aggregate findings and methodology improvements derived from Trial-tier use (see the Governor Trial Terms §4), but we do not sell vendor-facing clinical AI tools, and we do not accept licensing arrangements from the vendors whose products we evaluate.
The methodology is publicly documented and pre-registered on OSF before analysis runs. Any credentialed researcher can reproduce a confirmed finding. The pre-registration discipline is the moat, not the model weights and not the codebase. See the methodology page for the long-form walk-through and osf.io/3ws8g for the public protocol.
The Evaluator is institutional infrastructure, not a clinician-facing app. Output is shaped to feed governance review processes, not individual case-by-case decision support.
The system focuses where the disagreement is. Concordant cases skip expensive analysis, which is what makes the unit economics work at hospital scale.
Every evaluation gets four things stamped at the moment it is produced: an evidence hash over inputs and outputs, an Ed25519 signature from the engine's private key, a timestamp, and a methodology version string identifying the exact rule set in force when the evaluation ran.
The signing key's public half is published at a stable URL and rotates on a documented schedule. Anyone holding a finding ID or evidence hash can verify a record's provenance without calling us. The verifier is open: rocsitediscovery.com/verify/. Paste a finding ID or hash, get back the signed record and the engine fingerprint that produced it.
Try it now. Paste a finding ID into the verifier and read what comes back.
The verify endpoint is the proof that the cryptographic claim on this page is real, not aspirational. It's been live since the pre-registration registry was built.
Access to the Evaluator is granted to qualified institutional evaluators on request. We work directly with AI Governance Committees, radiology department leadership, and institutional procurement teams to scope evaluations. There is no self-serve signup. Conversations start with what you're trying to evaluate and why.
Tell us briefly who you are and what you'd like to evaluate. Adam will reply directly.
The Evaluator does not currently hold FDA clearance and is not represented as a clinical decision-support device. It is institutional governance infrastructure, intended to support committee review, not to replace clinical judgment on individual cases.