Methodology

How we rate clinical AI

We rate each tool in five categories using a published rubric applied to verified, sourced facts about the tool, not a hidden benchmark. Every rating traces to a fact you can check, and we publish every change.

What these ratings are, and what they are not

Each tool is rated from 0 to 10 in five categories. These are editorial ratings derived from verified facts, not measured accuracy scores. We do not run a proprietary accuracy benchmark, and we do not publish a number we cannot trace back to a documented, checkable fact about the tool, its pricing, availability, content sources, features, integrations, and governance.

Where independent, published studies of a tool’s accuracy or safety exist (for example peer-reviewed comparisons, or the Stanford/Harvard/ARISE NOHARM safety study), we cite and attribute them, we never restate a vendor’s marketing accuracy claim as our own finding.

The overall score is the simple average of the five categories. The choice of categories and their equal weighting is an editorial judgment, stated openly here so you can decide whether you agree with it. A reader who weights one category more heavily than we do may reasonably rank the tools differently.

The five categories

01, Accessibility

Who can actually use the tool. We weigh cost (free, freemium, paid, or institutional), geographic reach, professional-verification gates, and whether the product is advertising-funded. A paid or institution-only tool, or one limited to a single region, scores lower here regardless of how good its content is.

02, Content & Evidence

The substance behind the answers. We weigh the breadth and curation of the underlying corpus, whether answers carry citations traceable to a primary source, and how transparently the strength of evidence is graded. We do not score raw accuracy here, that would require a benchmark we do not run.

03, Clinical features

The breadth of decision-support functions a clinician can actually use: differential diagnosis, drug and interaction reference, treatment and management guidance, documentation/scribing, and exam preparation. More relevant, verifiable features score higher.

04, Integration

How well the tool fits a real workflow: EHR integration (e.g. SMART on FHIR), a developer API, and availability across web and mobile. A capable but standalone tool scores lower here than one embedded where clinicians already work.

05, Trust & governance

The safeguards around the tool: HIPAA and GDPR posture, independent validation, regulatory status, and conflicts of interest (for example, an advertising or pharmaceutical funding model). Independent third-party recognition raises this score; unaudited vendor claims do not.

What we don’t do

  • We don’t accept payment, free credits, or beta access in exchange for coverage.
  • We don’t let vendors review ratings before publication.
  • We don’t publish a score we can’t trace to a documented fact, and we don’t present editorial ratings as measured benchmark results.
  • We don’t change ratings retroactively without publishing what changed and why.

Conflicts and disclosures

Where any editorial team member has a prior relationship with a vendor we cover, that relationship is disclosed on our disclosures page and the member is recused from rating that tool. Where a tool is referenced as a “top pick” or “highlight, ” that placement must be defensible from the published facts and rubric, never editorial preference alone.

Updates and corrections

Ratings are reviewed quarterly, and a rating changes only when a verifiable fact about the tool, pricing, availability, features, governance, actually changes, with the change logged and dated. If we get a fact wrong, we publish a correction on the affected page and update our editorial policy log.