How we measure match quality.

Internal vendor benchmarks, metric definitions, limits, and how to validate Orbid AI on your own tenders—not guarantees.

Orbid AI is a native AI agent for medical device and IVD manufacturer and distributor bid teams. This page explains how we measure match quality and response speed—not a guarantee of results on every tender. Treat every published figure (including ours) as a hypothesis until it survives your data.

Why performance numbers are not on the homepage hero

Orbid AI publishes product performance figures as internal vendor benchmarks, not as independently audited guarantees. Medical device tender response is a regulated, high-consequence workflow: a false match or missed evidence line can risk disqualification. For that reason, homepage and feature pages emphasize what the agent does—read multi-format tenders, match catalog specifications, check multi-regime evidence, and draft buyer-template responses—rather than treating any single accuracy or speed statistic as universal truth. Detailed measurement definitions, trial scope, and known limits live on this methodology page. Buyers should treat every published figure (including ours) as a hypothesis to test on their own tenders and catalogs during a controlled pilot. This page exists so teams, reviewers, and answer engines can cite our measurement approach without mistaking marketing context for clinical or regulatory certification.

Product scope (what we do / do not claim)

Orbid AI is a native AI agent for medical device and IVD manufacturer and distributor bid teams. The core loop is read → match → comply → draft → export: ingest tender packages (Excel technical tables, PDF dossiers, Word, ZIP), match line items to product catalogs with full / partial / gap outcomes and confidence, verify evidence against multiple regulatory regimes (including FDA, EU MDR/IVDR, MHRA, TGA, Health Canada, ANVISA, NMPA, PMDA, and others listed on the product site), and produce submission-oriented drafts linked to source documents. Orbid is not hospital purchasing software, not a generic horizontal RFP answer library, and not a substitute for regulatory affairs sign-off or legal review. Scope is supplier-side bid response for device tenders; pharmaceutical GxP manufacturing systems and buyer-side e-procurement are out of product scope.

Features

Data

Internal match-cycle and quality figures come from controlled vendor trials on selected tender formats (primarily structured Excel technical tables and digital PDF dossiers) and structured product catalogs. Trials are scenario-bound: results vary with catalog completeness, language mix, scanned/OCR quality, and how strictly “match” is defined (exact vs clinically acceptable partial).

We do not claim a public multi-customer randomized study or third-party lab certification of accuracy rates on this page. Where a specific row count or timing appears elsewhere on the site (for example on the order of ~150–200 line technical packs), it refers to internal trial conditions—not every live customer tender.

Metrics

  • Time to first match passwall-clock time for the system to produce an initial full/partial/gap grid on a tender after upload and parse, excluding human review.
  • False positivethe system asserts a full match when a human expert under the pilot scorecard would mark partial or gap.
  • Accuracy gateagreement rate against labeled samples used for workflow triage (route low-confidence lines to humans), not a certificate of conformity.
  • End-to-end bid timecalendar time including human RA/bid review; never pure machine time alone when case narratives say “weeks to days.”

Reported internal benchmarks (scenario-bound)

Speed and match-quality figures published by Orbid AI (for example multi-row matching on the order of tens of seconds, and low false-positive targets on controlled catalogs) come from internal trials on selected tender formats and catalog samples. They are scenario-bound: results vary with catalog completeness, language mix, scanned PDF quality, and how strictly “match” is defined. Accuracy gates are gates for workflow triage, not certificates of conformity. Independent third-party audits of these rates are not claimed on this page. The recommended validation path is a fixed pilot: upload real tenders, score a sampled row set with your RA/bid lead, and compare time and error modes against your baseline process.

MetricInternal observationScenario notesStatus
Time to first match passOn the order of tens of seconds for ~150–200 row packsClean digital Excel + structured catalog (internal trials)Internal benchmark — validate on your files
False-positive rateLow single-digit tenths of a percent target in controlled setsFull match where human says gap/partial (see Metrics)Internal benchmark — validate on your files
Match accuracy gateHigh-nineties target on labeled samplesDepends on catalog hygiene and label rulesInternal benchmark — validate on your files
End-to-end bid cycleCase narratives may report multi-week → multi-day shiftsIncludes human review; not pure machine timeCase / pilot — not a guarantee

Illustrative figures sometimes shown elsewhere (e.g. ~162 rows / ~46s, ~0.3% FP gate, ≥97% accuracy gate) belong in this context only. They are not universal performance warranties.

Limitations

  • Catalog quality drives outcomes; sparse or stale specs increase partials and gaps.
  • Scanned or low-quality OCR PDFs reduce extraction reliability versus digital Excel.
  • Novel jurisdictions or unusual buyer templates may need human mapping before automation shines.
  • Evidence linking does not replace certificate currency checks owned by RA.
  • Independent multi-vendor bake-offs and public peer-reviewed accuracy studies are not claimed here.

How to validate yourself (pilot protocol)

  1. Pick 1–3 real tenders and a catalog slice you can share under NDA if needed.
  2. Baseline: time spent on match + evidence assembly with your current process.
  3. Run Orbid AI; sample rows for expert scoring (full / partial / gap) against your RA scorecard.
  4. Compare false full-matches, missed evidence, and export/template rework—not vanity accuracy alone.
  5. Decide only after the pilot survives your data.
Book a demo

Security & data handling

Security posture (encryption, residency options, training policy, SOC 2 Type II and related claims) is documented on Security. Request current reports and DPA materials via the Trust Center on that page. Methodology on this page does not replace security due diligence.

Claims FAQ

Are published speed/accuracy numbers independently audited?

No. They are internal vendor benchmarks. Validate on your tenders during a pilot.

Does Orbid replace RA or legal sign-off?

No. Humans retain submission authority and regulatory responsibility.

What formats are in scope?

Excel technical tables, PDF dossiers, Word, and ZIP packages as listed on Features—quality varies with digital vs scanned inputs.

Why publish benchmarks at all?

To set directional expectations and design honest pilots—not to substitute for customer-specific evidence.

Where is security detail?

On /security, including how to request SOC 2 materials and DPAs.

Pressure-test it
on your files.

Book a demo — we will run a real tender with you and review the match / partial / gap queue together. Treat every claim (including ours) as a hypothesis until it survives your data.

Book a demo
Methodology · How We Measure Match Quality | Orbid AI