Built by SpreadRun

Clinical Trial Results Table QA

Validate normalized clinical-trial study and outcome tables before they feed your analysis, your joins, or your customers. Submit your tables; get a deterministic audit that flags malformed NCT IDs, missing required fields, duplicate records, orphan outcomes, invalid outcome types and bad results-posting dates, with exact row locations for every finding. $0.25 per completed audit. No subscription.

What it checks

One audit runs every rule, deterministically, every time:

  • NCT ID validity. Every trial_id in both tables must be NCT followed by exactly 8 digits. Malformed IDs are listed with their row.
  • Required fields. The required columns must exist, and every row must have a value in each of them. See the required-column list.
  • Duplicate records. One row per trial_id in the studies table, one row per trial_id plus outcome_id in the outcomes table. Exact repeats are flagged.
  • Orphan outcomes. Outcome rows that reference a trial_id with no matching study row: the join-breaker.
  • Outcome type. Must be PRIMARY, SECONDARY or OTHER_PRE_SPECIFIED.
  • Results posting date. results_first_post_date must be a real calendar date written as YYYY-MM-DD, so 2026-02-30 is caught.
  • Structural integrity. Unique headers, the same number of cells in every row, well-formed CSV. Tables that fail these are rejected before the audit and not charged.

Every finding includes the table, row number, field, and the rule that fired, so fixing is mechanical, not detective work. The report also gives the share of non-empty values per required column.

Who it's for

  • Researchers and evidence-synthesis teams: audit extracted trial tables before meta-analysis or systematic review. A bad join upstream poisons everything downstream.
  • Data vendors: QA the trial datasets you sell or license, as a gate before each delivery.
  • Developers: validate tables pulled from the ClinicalTrials.gov API v2 or the AACT database after your own normalization.
  • AI agents and automation: REST endpoint, JSON in and out, deterministic results a downstream agent can act on without a person in the loop.

Limits per audit: 10,000 rows and 40 columns per table, request body up to 4.4 MB, the first 1,000 findings listed in full with complete counts.

How it works

  1. Submit

    Send your normalized studies table and outcomes table as CSV text in one JSON request, or upload them in the test form below.

  2. Audit

    Every check runs deterministically against the documented rule set. Same input, same report, every time.

  3. Fix from the report

    Each finding comes with its exact location. Correct your source data and re-run to confirm clean. A clean re-run is your QA evidence.

Try it now

Free demo, no account: up to 512 KB per run and 10 runs a day. For full validations from this form, sign in and buy credits: $0.25 per completed report, packs from $5.

Paste CSV, upload files, or load a synthetic sample.

Columns: trial_id, condition, phase. One row per trial.
Columns: trial_id, outcome_id, outcome_type, primary_endpoint, result_status, results_first_post_date.
Sample report
Synthetic tables with six kinds of defect
FAIL
6 studies7 outcomes7 findings
TableRowFieldRule
studies4trial_idINVALID_NCT_ID
studies6trial_idDUPLICATE_KEY
outcomes4primary_endpointMISSING_VALUE
outcomes5outcome_idDUPLICATE_KEY
outcomes3results_first_post_dateINVALID_ISO_DATE
outcomes6trial_idORPHAN_OUTCOME
outcomes7outcome_typeINVALID_OUTCOME_TYPE

Pricing

$0.25 per completed audit. One submission (studies table plus outcomes table), one full report.

  • Billed only when an audit completes and produces a report, PASS or FAIL. Unreadable or malformed inputs are not billed.
  • No subscription, no seats. Prepaid credits from $5 (20 audits). Credits never expire. All pricing
  • Running it on a schedule? Each call is one audit, so call it from your scheduler or CI. Questions about volume: contact us.

Call it from code

Full request and response reference, error codes and limits are in the API docs.

curl -X POST https://www.spreadrun.com/api/v1/clinical-trial-table-validator \
  -H "Authorization: Bearer $SPREADRUN_API_KEY" \
  -H "Content-Type: application/json" \
  --data "$(jq -n --rawfile s studies.csv --rawfile o outcomes.csv \
            '{studiesCsv: $s, outcomesCsv: $o}')"

Questions

What exactly do I submit?

Two CSV tables, sent as CSV text inside one JSON request. The studies table needs the columns trial_id, condition and phase, one row per trial. The outcomes table needs trial_id, outcome_id, outcome_type, primary_endpoint, result_status and results_first_post_date, one row per outcome measure, linked to studies by trial_id. Other columns are allowed (up to 40 per table) and ignored.

Where does the source data usually come from?

Most users pull from the ClinicalTrials.gov API v2 or the AACT (Aggregate Analysis of ClinicalTrials.gov) database, then normalize into their own schema. The audit validates your normalized tables, whatever the upstream source.

Does this replace ClinicalTrials.gov's own quality control?

No. ClinicalTrials.gov's PRS system validates records submitted to the registry. This validates the tables you extracted and normalized for your own analysis or product: a different stage, with different failure modes (bad joins, dropped rows, mangled dates from your ETL, not theirs).

Will it fix my data?

No, deliberately. It reports, with exact locations. Automated fixing of trial data would be irresponsible; a person, or your pipeline logic, decides each correction. Re-run after fixing to confirm clean.

What's an "orphan outcome" and why does it matter?

An outcome row whose trial_id matches no row in the studies table. It usually means a dropped or mismatched study record upstream, and it silently breaks any join between the two tables. It is one of the most common silent-corruption patterns in merged trial datasets.

How is pricing counted?

$0.25 per completed audit: one submission of your table pair, one full report, PASS or FAIL. Requests rejected before an audit runs (bad JSON, missing columns, ragged rows, over the size limit) are free. Re-runs after fixes are billed the same way, which is why you fix from the report rather than guessing.

Is my data stored or shared?

Your tables are processed in memory for the length of the request and are not stored or shared. The report does not repeat your cell values. For billing and usage we log the time, endpoint, result status, request size and duration, never the table contents.

Can I use this inside an automated pipeline?

Yes, that is the primary design case. REST endpoint, API key, structured JSON report with machine-readable finding codes. A downstream step can fail the build whenever status is FAIL.

Does it check clinical accuracy or regulatory compliance?

No. It is structural QA of your tables. It does not judge whether results are clinically correct, current, authentic, or compliant with reporting rules.

More validators: Hospital Price Transparency MRF Validator | the full catalog