Back to overview

Dataset v5 · Evaluated Jul 26, 2026

How the Trust Score works.

The score is an editorial model applied to a manually researched, source-backed corpus. Its rules, limits, source package, and migration trail remain as inspectable as its conclusion.

01

Scope

What this dataset is—and is not.

The corpus contains 155 concrete Elon Musk statements, commitments, forecasts, promises, and factual assertions spanning Business & Technology, Media & Society, Personal & Legal, Politics & Government, and Science & Health. Company affiliation is not required: the same subject-first taxonomy covers organizational claims, public affairs, and statements about science, media, legal matters, or Musk personally.

It is not literally every sentence Musk has spoken or posted. That universe is unbounded and includes deleted material, private remarks, jokes, opinions, and claims with no testable meaning. The corpus deliberately favors consequential, concrete, verifiable claims. The optional Topic category is populated on 47 records across 8 cross-cutting topics: Accusations About Individuals, Government & Public Spending, Immigration & Demographics, Legal & Regulatory, Media & Information, Personal & Biographical, Politics & Elections, and Science & Health.

The central limitation is selection bias.

This is not a statistically random sample. The public-discourse topic-tagged subset is especially adverse-selected because fact-checkers examine disputed statements rather than random everyday remarks. The score must not be described as the percentage of everything Musk says that is true.

02

Sources

How evidence is prioritized.

  1. Original record.Musk’s statement, company material, filing, court or regulator record, or another official source.
  2. Major reporting.Reuters, Associated Press, The New York Times, or another major outlet documenting the statement and measurable outcome.
  3. Transparent fact-checking.Established specialist fact-checkers when they show their sources and link the original claim.
  4. Secondary material.Used only when stronger evidence is unavailable, with research confidence reduced.

Every row includes a statement-source URL and at least one outcome-source URL. The current package contains 392 citation placements across 200 distinct URLs. Statements are paraphrased to reduce quotation errors and preserve fair-use restraint; the linked source carries the original context.

Evidence structure is scored and audited separately.

Each of the 155 claims has one source-audit record. The component scores describe how strong and direct the cited evidence is; they are not decimal-level probabilities that a verdict is correct.

30

Statement evidence quality

Quality and directness of evidence that Musk made the statement.

40

Outcome evidence quality

Quality of evidence establishing the outcome.

15

Corroboration

Independent corroboration across sources and domains.

15

Directness

How directly the evidence resolves the proposition.

The four components sum to 100. Verdict confidence begins with evidence strength and applies the published deductions for nuance, credible-source disagreement, unresolved conflicts, interested-party evidence, or missing independent corroboration. The current strict-promise-evidence-audit-v2.0 package contains 155 source audits and 1860 field-level evaluation audits covering 12 outputs per claim.

03

Scoring

5 scored categories, 2 visible exclusions.

The v5 Trust Score is total points earned divided by total points possible. The current corpus earns 3,100 of 13,800 possible points across 138 scored claims: 22.5%, rounded to 22% for the homepage.

100

True

Promise/forecast verdict: Every material term was met, including timing, scope, quantity, capability, price, and permanence. Factual verdict: The central factual proposition is supported by strong evidence.

75

Mostly True

The core proposition is accurate but has a meaningful nonfatal qualification.

50

Misleading

Some facts are accurate, but omitted context or framing materially changes the reasonable takeaway.

25

Unsupported

An affirmative claim lacks adequate credible support but is not conclusively disproven.

0

False

Promise/forecast verdict: At least one matured material term failed, including a deadline miss, incomplete scope, reversal, or incorrect forecast. Factual verdict: Reliable evidence contradicts the central proposition.

Pending and Unresolved rows receive no score and remain visible. Every score-bearing category and point value above comes directly from the current classification-key CSV.

Promises and forecasts use a strict binary test.

A matured promise or forecast passes only when every material term—including timing—is met. A missed deadline fails the original proposition even if the promised result arrives later. Pending and genuinely unresolved records stay visible but are excluded until they can be resolved.

04

Verdicts

Canonical categories and contextual display labels.

True

Promise/forecast verdict: Every material term was met, including timing, scope, quantity, capability, price, and permanence. Factual verdict: The central factual proposition is supported by strong evidence.

Mostly True

The core proposition is accurate but has a meaningful nonfatal qualification.

Misleading

Some facts are accurate, but omitted context or framing materially changes the reasonable takeaway.

Unsupported

An affirmative claim lacks adequate credible support but is not conclusively disproven.

False

Promise/forecast verdict: At least one matured material term failed, including a deadline miss, incomplete scope, reversal, or incorrect forecast. Factual verdict: Reliable evidence contradicts the central proposition.

Pending

The stated deadline or condition had not matured by the evaluation date.

Unresolved

Promise/forecast verdict: Available evidence cannot responsibly establish pass or fail. Factual verdict: Evidence is insufficient or materially conflicted.

Claim type determines the available verdicts.

Promises and forecasts use True or False once resolved under the strict material-terms rule. Mostly True, Misleading, and Unsupported remain available only for factual assertions, where evidence can support a nuanced finding.

Contestation is a badge, not a verdict.

57 records are explicitly contested by credible sources. Contestation remains independent of the primary verdict after the total evidence is evaluated.

A lie requires evidence of intent.

False, late, unsupported, or reversed does not automatically mean deliberate deception. Every row provides a public Yes, No, or Not assessable answer and a more detailed intent status, while avoiding an inference about state of mind from outcome alone.

Was intentional deception established?

Public answer

Yes
0
No
138
Not assessable
17

Detailed assessment

Established
0
Suggested but not established
5
Not established
133
Not assessable
17

“No” means the cited evidence does not establish intentional deception. It does not convert a False, Misleading, Unsupported, or unfulfilled claim into a true one. “Not assessable” is reserved for Pending or Unresolved records.

05

Rating bands

How the conclusion is assigned.

80–100Highly Trustworthy
65–79.9Generally Trustworthy
45–64.9Inconsistent
25–44.9Not Trustworthy
0–24.9Highly Untrustworthy

These thresholds are editorial judgments, not a scientific standard. At 22.5, the current corpus falls within “Highly Untrustworthy.”

06

Structure

Subject categories are primary; entities remain separate.

  1. Personal & Legal.9 tracked records; 8 included in the score.
  2. Media & Society.17 tracked records; 15 included in the score.
  3. Business & Technology.92 tracked records; 79 included in the score.
  4. Politics & Government.29 tracked records; 29 included in the score.
  5. Science & Health.8 tracked records; 7 included in the score.

Every row has one of these 5 subject categories in primary_domain. The organization_or_domain field is an independent context facet, currently containing 9 distinct values; it does not determine what the claim is about. The optional public_discourse_category field adds a cross-cutting Topic category wherever it applies, including claims associated with a company. relationship_to_organization describes how the organization or context relates to Musk.

Claim type describes the statement's form, not its subject. The current statement forms are:

  1. Promise or Commitment.51 tracked records; 48 included in the score.
  2. Factual Assertion.57 tracked records; 55 included in the score.
  3. Prediction or Forecast.46 tracked records; 35 included in the score.
  4. Opinion or Rhetoric.1 tracked record; 0 included in the score.

All names and counts above come from the row-level CSV and update automatically when the data package changes. None of these context fields limits the subject taxonomy to organizations Musk has owned.

07

External context

A broader benchmark, kept separate.

A June 2026 New York Times audit reviewed more than 69,000 Musk social posts and 19 Tesla investor calls, then manually evaluated 602 concrete future-deadline business goals. It reported 19% achieved on time, 35% late or unfulfilled, 33% unclear or without a public update, and 13% with future deadlines.

That research is directionally relevant but uses different selection and scoring rules. Its entries are not represented as rows in this corpus and do not affect this Trust Score.

Read the New York Times audit (opens in a new tab)
08

Maintenance

Corrections should leave a trail.

  • Preserve every existing record ID and correction history.
  • Add new claims as new rows and document corrections in migration.
  • Recheck Pending and Unresolved rows on a fixed cadence.
  • Keep evaluation dates and outcome sources current.
  • Publish corrections and scoring changes in version control.
  • Require a second review for high-impact, personal, legal, medical, political, or intent-sensitive classifications.
  • Recompute the summary directly from the row-level CSV.

Correction and repetition fields describe what is documented in each row's cited evidence. They are not presented as an exhaustive search of every deleted post, interview, reply, or later repetition.

v5

92 records reclassified; 51 records audited, no headline change; 12 records new v5 record. The row-by-row migration CSV preserves the complete update trail.

09

Downloads

Inspect or reuse the complete v5 package.