Five years of news, read from thousands of perspectives
Thinking Text, a partner company, analysed 15 million news articles covering five years — close to 1 trillion tokens processed across multiple research iterations — to build daily oil and equity indexes. The engine behind it is spinf.
news articles
tokens processed
correlation with crude
scores a day, oil alone
Fixed questions, asked of every article, every day
The indexes do not ask a model for an opinion on the market. They ask simple, fixed questions of every article and let scale do the work: the answers are aggregated into a read of what the press says, and only then compared with prices.
Collect
Every English article from about two hundred selected outlets, public and private feeds, de-duplicated. Roughly a thousand oil articles a day.
Score
spinf reads each article once and answers about five hundred questions: every topic asked through several formulations. No price language, no event extraction.
Calibrate & aggregate
Scores are measured against their empty floor and a baseline of articles, combined over the day and standardised on each question’s own trailing year.
Relate
The daily read is compared with WTI and Brent: correlation with the price level, the news-implied move at 1, 5 and 20 days, and what is left unexplained.
From a template to scored cells
The heart of the method is the cell: one article, one question, one probability. Topics are written as templates over a five-step scale, expanded into questions, and scored together in one call per article.
Global oil demand is {w}.
Each word of the scale makes one question. Every topic is also asked through several formulations — “Global oil demand is …”, “Worldwide consumption of crude oil is …”, “Demand for oil and fuels is …” — so one topic becomes 15+ questions, all answered from a single read of the article.
Illustrative values. Shaded cells sit above their empty floor. Score = position on the −2…+2 scale weighted by each question’s lift over the floor; weight = total lift. Article C barely talks about demand: its lifts are small and its residual is high, so it hardly moves the day’s demand read. Simplified — the production aggregation is private.
Raw probabilities become absolute ones against a baseline
A raw p depends on how the question is worded. spinf returns what you need to remove that bias: subtract the empty floor, then place the score on the distribution of the same question over a baseline of comparable articles. The result is a probability you can compare across questions, sources and years.
Floor: remove what the question scores with no content at all.
Residual: down-weight questions that do not apply to the article.
Baseline: rank each score against a reference set of articles.
Aggregate: combine thousands of cells per day, standardise per question.
One raw score, placed on its baseline
Article A scores p = 0.71 on this question, higher than 93.8% of 10,000 baseline articles (z ≈ 1.8).
Synthetic illustration. The highlighted bar holds article A. Converting raw p into a rank or a z against a baseline of the same kind of content turns it into an absolute, comparable probability.
A read of the press that tracks the price of crude
The oil read (OIL-F6) against WTI and Brent since mid-2022. The level correlation tops 0.8 in three of the last four years, the move implied by the news points the right way on 92% of big 20-day moves, and six simple questions beat keyword counting on the same articles by a wide margin.
OIL-F6 read vs WTI (CL)
Since mid-2022 · correlation 0.80 on 1064 market days
The spinf read (six fundamental questions, averaged over ~1,000 articles a day) against the log price minus its trailing 60-day median. Both series scaled by their own standard deviation. Data: Thinking Text OIL-F6, live.
Correlation with the WTI price level, by year
Pooled 0.80 · 2022–2026
Read level vs log price minus its trailing 60-day median. 2026 is year to date.
Questions vs keyword counting, same articles
2022-07-01 → 2026-09-02
Correlation with the price level (log price minus its 60-day median).
Every day, a detailed and inspectable read
Each of the six questions is reported with its standardised value, its state and the number of scores behind it — and the day's crude move is split into what the news implies and what it leaves unexplained.
Full methodology notes and downloads on thinkingtext.com.
Oil News Index F6 · score card
Press day 2026-09-26 · final
+0.10
read (z) · balanced531
Articles read90
Outlets829
Distinct prompts54,343
Answers (scores)WTI, 5-day move
−7.4%
actual−2.3%
news-implied−5.1%
residualBack-analysis: sign right on 88% of big moves · R² 0.32 out of sample
WTI, 20-day move
+14.0%
actual+2.9%
news-implied+11.1%
residualBack-analysis: sign right on 92% of big moves · R² 0.37 out of sample
Company reads for Nasdaq-100 names
The equity index runs on the same engine, riding on the oil read for about a minute of GPU a day: company, peer and market questions on every article that names the company.
Equity News Index · published names
Correlation of the weekly change of each company read with the weekly stock return, since 2022.
Data: Thinking Text EQ-S10, live. Articles per day on the latest press day.
What would you ask a million documents?
News is one example. The same engine codes surveys and reviews, triages tickets and moderates posts: anywhere many questions meet many documents.