Success story · Thinking Text news indexes

Five years of news, read from thousands of perspectives

Thinking Text, a partner company, analysed 15 million news articles covering five years — close to 1 trillion tokens processed across multiple research iterations — to build daily oil and equity indexes. The engine behind it is spinf.

15M
news articles
five years, 2022–2026
~1T
tokens processed
across research iterations
0.85
correlation with crude
best year · 0.80 pooled on WTI since 2022
54,343
scores a day, oil alone
531 articles × 829 prompts · 2026-09-26
The approach

Fixed questions, asked of every article, every day

The indexes do not ask a model for an opinion on the market. They ask simple, fixed questions of every article and let scale do the work: the answers are aggregated into a read of what the press says, and only then compared with prices.

Collect

Every English article from about two hundred selected outlets, public and private feeds, de-duplicated. Roughly a thousand oil articles a day.

Score

spinf reads each article once and answers about five hundred questions: every topic asked through several formulations. No price language, no event extraction.

Calibrate & aggregate

Scores are measured against their empty floor and a baseline of articles, combined over the day and standardised on each question’s own trailing year.

Relate

The daily read is compared with WTI and Brent: correlation with the price level, the news-implied move at 1, 5 and 20 days, and what is left unexplained.

Base scoring

From a template to scored cells

The heart of the method is the cell: one article, one question, one probability. Topics are written as templates over a five-step scale, expanded into questions, and scored together in one call per article.

1 · Template

Global oil demand is {w}.

falling sharply (-2)
falling (-1)
steady (0)
rising (+1)
rising sharply (+2)

Each word of the scale makes one question. Every topic is also asked through several formulations — “Global oil demand is …”, “Worldwide consumption of crude oil is …”, “Demand for oil and fuels is …” — so one topic becomes 15+ questions, all answered from a single read of the article.

2 · Cells: article × question → p
Articlefalling sharplyfallingsteadyrisingrising sharplyresidualscoreweight
Empty floor (no content)0.060.140.300.180.07
AChina’s crude imports hit a record as refiners restock ahead of winter.0.020.040.120.710.380.02+1.370.84
BIEA trims its demand growth forecast, citing weaker industrial activity in Europe.0.080.660.310.090.020.03-1.020.55
CHurricane forces evacuation of Gulf of Mexico platforms as shut-ins mount.0.050.120.330.200.060.29+0.400.05

Illustrative values. Shaded cells sit above their empty floor. Score = position on the −2…+2 scale weighted by each question’s lift over the floor; weight = total lift. Article C barely talks about demand: its lifts are small and its residual is high, so it hardly moves the day’s demand read. Simplified — the production aggregation is private.

Calibration

Raw probabilities become absolute ones against a baseline

A raw p depends on how the question is worded. spinf returns what you need to remove that bias: subtract the empty floor, then place the score on the distribution of the same question over a baseline of comparable articles. The result is a probability you can compare across questions, sources and years.

  1. Floor: remove what the question scores with no content at all.

  2. Residual: down-weight questions that do not apply to the article.

  3. Baseline: rank each score against a reference set of articles.

  4. Aggregate: combine thousands of cells per day, standardise per question.

One raw score, placed on its baseline

Article A scores p = 0.71 on this question, higher than 93.8% of 10,000 baseline articles (z ≈ 1.8).

Synthetic illustration. The highlighted bar holds article A. Converting raw p into a rank or a z against a baseline of the same kind of content turns it into an absolute, comparable probability.

Results · live data

A read of the press that tracks the price of crude

The oil read (OIL-F6) against WTI and Brent since mid-2022. The level correlation tops 0.8 in three of the last four years, the move implied by the news points the right way on 92% of big 20-day moves, and six simple questions beat keyword counting on the same articles by a wide margin.

OIL-F6 read vs WTI (CL)

Since mid-2022 · correlation 0.80 on 1064 market days

View
Price
Year

The spinf read (six fundamental questions, averaged over ~1,000 articles a day) against the log price minus its trailing 60-day median. Both series scaled by their own standard deviation. Data: Thinking Text OIL-F6, live.

Correlation with the WTI price level, by year

Pooled 0.80 · 2022–2026

Read level vs log price minus its trailing 60-day median. 2026 is year to date.

Questions vs keyword counting, same articles

2022-07-01 → 2026-09-02

Correlation with the price level (log price minus its 60-day median).

Score card

Every day, a detailed and inspectable read

Each of the six questions is reported with its standardised value, its state and the number of scores behind it — and the day's crude move is split into what the news implies and what it leaves unexplained.

Full methodology notes and downloads on thinkingtext.com.

Oil News Index F6 · score card

Press day 2026-09-26 · final

+0.10

read (z) · balanced

531

Articles read

90

Outlets

829

Distinct prompts

54,343

Answers (scores)
QuestionToday's readingzStateAnswers
Global oil demandsteady+0.08842
Global oil supplysteady+0.011,091
US oil inventoriessteady−0.101,597
OPEC+ production quotastightening+0.5386
Supply disruption risksteady−0.05865
Refining marginssteady−0.02617

WTI, 5-day move

−7.4%

actual

−2.3%

news-implied

−5.1%

residual

Back-analysis: sign right on 88% of big moves · R² 0.32 out of sample

WTI, 20-day move

+14.0%

actual

+2.9%

news-implied

+11.1%

residual

Back-analysis: sign right on 92% of big moves · R² 0.37 out of sample

Same engine · equities

Company reads for Nasdaq-100 names

The equity index runs on the same engine, riding on the oil read for about a minute of GPU a day: company, peer and market questions on every article that names the company.

Equity News Index · published names

Correlation of the weekly change of each company read with the weekly stock return, since 2022.

Company5-day corr.worst yeararticles / day
Tesla TSLA
0.640.6041
AMD AMD
0.580.5224
Nvidia NVDA
0.570.5668
Intel INTC
0.570.4620
Palantir PLTR
0.560.447
Micron MU
0.560.3413
Meta META
0.530.5349
Amazon AMZN
0.530.3779
Alphabet GOOGL
0.500.40130
Apple AAPL
0.480.3183
Microsoft MSFT
0.470.3884
Cisco CSCO
0.410.355

Data: Thinking Text EQ-S10, live. Articles per day on the latest press day.

Your use case

What would you ask a million documents?

News is one example. The same engine codes surveys and reviews, triages tickets and moderates posts: anywhere many questions meet many documents.