This site maps the structure of financial news in New York across four decades of the Wall Street Journal, from the paper’s founding in 1889 to 1930. It is the fourth in a series after London, Shanghai and St Petersburg, and it uses the same window as the London site so the two financial centres can be read against one another over the gold-standard era, the Panic of 1907, the war, the 1920s boom and the Crash.
Two things are on offer. The topic model asks what the paper was about — how attention divided between the daily market report and the call-money market, cotton and grain, iron and steel, four separate railroad regions, bank clearings and Washington, and how that mix moved across the period. The Historical Recall section asks something the other sites cannot: which episodes of the past the paper reached for, and when.
Corpus. A ProQuest TDM Studio export of 2,225,857 Wall Street Journal records, 1889–1933. Restricted to 1889–1930 that is 1,962,436 records, of which 1,106,647 are running prose of at least 200 characters once stock-quote tables, display and classified advertising, tables of contents and front matter are set aside. Those 1.1 million articles are the topic-model corpus. Median OCR quality is 81%.
Topic model. Non-negative matrix factorisation over tf–idf, 50 topics, fixed seed. Vocabulary is filtered against an English dictionary plus a hand-kept list of period market terms, which removes the OCR fragments that otherwise dominate a corpus of this vintage. Topic labels are hand-written from the top-weighted terms.
Market-news share. The Journal is a financial paper end to end, so “what share of this topic is financial” is not the useful split. The figure reported instead is the share of a topic’s mass coming from articles where “stock” falls within 200 characters of “market” — the same proximity filter used on the published 1984–2020 WSJ corpus. It separates market news from corporate, commodity and general business reporting. On this corpus 71,754 articles qualify, 6.5% of the prose.
A note on what is measured. Attention here is column inches, not importance. A topic that rises is one the paper printed more of; nothing in the method speaks to whether it mattered.
Why there is no war topic. A reasonable thing to expect of a corpus spanning 1914–18, and its absence is a finding rather than a failure. War language is certainly present: it appears in 28.5% of articles in 1918 against under 3% in the 1890s. One topic does carry it — War, government & foreign affairs rises from 0.9% of the paper in 1889 to 4.5% in 1918 — but it reads as a blend rather than a war topic.
Refitting on 1914–19 alone, when the war dominated the news, still yields only one war topic in twenty; the other nineteen remain cotton, copper, steel, grain, clearings and dividends. Tracing individual words explains why. Political vocabulary concentrates almost entirely in that one topic — enemy 99%, army 100%, treaty 97%. Financial war vocabulary scatters: liberty, as in Liberty Bonds, puts only 33% of its weight in any single topic, spreading across banks and reserves, bond grades and government; victory, as in the Victory Loan, 42%; munitions 58%, divided between steel and copper.
The paper covered the war twice over. The political war became a subject. The financial war did not — it entered as bond issues, industrial orders and tax measures, filed under banking, steel and copper. For a financial daily the war was less a topic than an adjective on the topics it already had.
Fifty topics over 1,106,647 prose articles. Each line is a topic’s mean share of articles in that year, so the series answers “how much of the paper was this?” rather than “how many articles were there?” — the paper roughly triples in size across the window, and a raw count would show that growth and little else.
Topics that rise and fall together, clustered on the correlation of their yearly shares. Co-movement is not subject matter: two topics cluster because the paper gave them space in the same years, which may mean they described one story, or merely that they competed for the same column inches.
Words keep their spelling and change their meaning. A separate embedding model is trained on each decade of the corpus, the four models are rotated into a common space, and a word’s movement is the cosine distance between its positions. What the company a word keeps in the 1890s and the 1920s reveals is often more legible than the number.
Why alignment is needed. Each decade’s model is trained independently, and independently trained embedding spaces differ by an arbitrary rotation. A raw cosine between two decades would measure that rotation, not any change in meaning. The models are therefore aligned by orthogonal Procrustes onto the 1900s before anything is compared — the method of Hamilton, Leskovec and Jurafsky (2016).
Neighbour filtering. These are fastText models, which build a word from its character n-grams. Left alone, a word’s nearest neighbours are its own spellings and the OCR wreckage around them — equity returns equities, equit, quity. Two filters are applied: a candidate sharing a four-character stem with the query is dropped, and candidates must appear in a 69,905-word English dictionary.
Read the 1890s column with care. Early OCR is markedly worse — 63% of tokens are dictionary words in 1889 against 85% from 1900 — and it shows in the neighbours. In the 1890s crash sits beside cash, brash and cask; ticker beside picker and dicker. These are phonetic and scanning cousins that survive a dictionary filter because they are real words. Part of what the shift measure registers for such words is therefore the corpus getting cleaner, not the language changing. The movements worth trusting are the ones where both ends are substantive — pool from market combination to oil field, curb from the kerbside market to a listed exchange, reserve from a specie proportion to the Federal Reserve.
Comparison with London. The models use the same hyperparameters as the FT models — skip-gram, 100 dimensions, window 5, ten epochs, minimum count ten, character n-grams of three to six — so the two corpora can be read together. Two caveats on that comparison. New York spans four decades against London’s twelve, so distances here are net movement over roughly forty years, not a century, and the magnitudes are not comparable in absolute terms — only the ranking is. And the New York tokenizer keeps two-letter words where London’s dropped them, a small difference in what fell inside each training window.
Not what the paper covered but how it wrote — whether it looked forward or back, hedged or asserted, explained or merely reported, and how often it named a person or quoted one. Each measure is computed per article and averaged by month.
Conflict, causal explanation, moral judgment, metaphor, suspense and direct quotation, each as occurrences per thousand words.
How often the paper named a year more than two years older than the issue it appeared in — a crude, purely mechanical count of looking backwards. The Historical Recall section measures the same instinct properly, by reading what the reference was to.
These metrics are computed by the same module that produced the London site’s, imported rather than reimplemented, so the two are directly comparable: identical word lists, identical denominators, and the same filters — articles under 50 words, or more than 15% digits, are excluded, which drops roughly a quarter of the corpus. Detection is exact whole-word matching against fixed lists, with no parser and no sentiment model, so a metric labelled “moral judgment” counts words like fraud and integrity; it does not understand them.
OCR quality is included as a metric because it is a confound, not a finding. It measures the share of tokens that are common function words, which collapses when scanning fails. Where it dips, treat movement in the other series with suspicion: the paper may not have changed, only the legibility of the scan.
The 1901 spike in Numbers / Article is mostly tables. That series is a raw count, not a rate — the one metric in the set not divided by article length — and it reads 4.50× its baseline in 1901. Normalized per word the same year is only 1.62×, so about two thirds of the apparent surge is articles getting longer rather than the paper printing denser figures. The mean article runs 3,917 characters in 1901 against roughly 1,500 either side of it.
What lies underneath is a real and short-lived editorial habit. Between about 1899 and 1903 the Journal ran long composite tabular compilations — gross-and-net-earnings tables covering every important railroad system, capital-and- dividend tables, monthly comparative statements — and the digitised archive stores each as a single article record. Articles opening with table language (“in the following table will be found…”) are 0.2–0.3% of the corpus in normal years but 3.5% in 1901, a twelvefold jump. They survive the digit filter because the surrounding prose dilutes the figures below the 15% threshold. Numbers / 1,000 Words and Article Length are plotted alongside so the two effects can be separated; the raw metric is kept unchanged because altering it would break comparability with London.
One word list travels badly. Using London’s lists unchanged is what makes the two series comparable, but it imports London’s vocabulary along with them. Personalization counts sir, lord and lady among its eleven terms, honorifics a New York paper had little occasion to print, and the measure duly runs at 1.96 per article here against 4.23 in the Financial Times over the same years. Much of that gap is the word list, not the journalism. The other measures track far more closely — forward-looking 0.575 against 0.572, hedging 0.493 against 0.499 — which is the better evidence that the port is faithful.
Which episodes of the past did the Journal reach for, and when? Every reference below was read out of an article by a language model and then verified against a verbatim quote from that article — a reference counts only if the passage that triggered it can be produced. Those passages govern the counts but are not reproduced here: the underlying text is licensed from ProQuest TDM Studio, so it stays on the research side. This section has no counterpart on the other sites in the series.
Select an episode to see how often the paper reached for it, year by year.
The criterion is an event a reader would call textbook-worthy, or one later recognised as significant. Five exclusions bind, and they do most of the work: securities named after events do not count (“Dawes bonds” is not a reference to the Dawes Plan); current conditions do not count (an article describing a slump underway is not recalling one); vague gestures to “past panics” do not count; routine corporate events do not count; and nothing inferred rather than read on the page counts.
Denominator. Intensity is references per million words of labeled text — the 71,754 market articles actually scanned, not the full 1.1 million. Raw counts would track how much the paper printed rather than how much it looked back.
Known defects, stated rather than smoothed. Of 21,613 grounded, registry-resolved references in the window, 574 (2.7%) place an event after the article that supposedly recalls it — a registry dating error, since no article can remember what has not happened — and are dropped here rather than displayed. A further 436 references could not be matched to a registry entry and are excluded, and 21 model responses were cut off at a length limit and are incomplete. That leaves 21,039 references on this page. The shaded years at the left rest on very little text — 1889 is 277 articles — so their rates are noise, not trend.
Why this differs from the paper. The published figures are computed on the shared set, the 927 episodes named by both the Journal and the London Financial Times, which is what makes the two comparable. This page uses the full WSJ registry instead, so the counts here are larger and the rates differ: the 1929 Crash runs at 30.1 references per million words over 1889–1930 on this page against 35.8 on the shared set. Both describe the same corpus; they differ in what they divide by.
Everything the page draws on, as it draws on it.
| File | Contents |
|---|---|
| wsj_topics.json | 16 topics: label, market-news share, top-weighted terms |
| wsj_theta.json | Mean topic share by year, 1889–1930 |
| wsj_correlation.json | Topic–topic correlation over yearly shares |
| wsj_taxonomy.json | Average-linkage clustering of the correlation matrix |
| wsj_recall.json | Historical references: intensity by year, top episodes, category mix (counts only — no article text) |
| meta.json | Corpus counts |
The news is one half of the record. These are the price series that cover the same market over the same years. Most are held by the Yale International Center for Finance, which asks that any use be credited: “Data made available through the International Center for Finance at Yale University.”
| Database | Coverage | Source |
|---|---|---|
| Old New York Stock Exchange Project Monthly prices for 671 individual securities, plus annual dividend data. Collected at the Yale ICF; corrected 2020, gaps filled from the Commercial and Financial Chronicle in 2021. |
1815–1925 | Yale ICF
— New York Stock Exchange: monthly prices, annual dividends 1825–1870,
the price-weighted index, security names, and the underlying newspaper transcriptions. Goetzmann, Ibbotson & Peng (2001), “A new historical database for the NYSE 1815 to 1925: Performance and predictability,” Journal of Financial Markets 4, 1–32, doi:10.1016/S1386-4181(00)00013-6.
Check which vintage you have. The series has been revised twice,
and each revision is documented:
2021-09-04 carries both rounds; one dated
2020-08-20 carries only the first and is still missing the 21 gap
months; anything earlier predates both.
|
| Price Quotations in Early United States Securities Markets Ten markets — nine US cities and London. |
1786–1862 | Sylla, Wilson & Wright, ICPSR 4053. doi:10.3886/ICPSR04053.v1 |
| McQuarrie bank and transportation stock indexes An independent reconstruction of early US equity returns, extending and in places disputing the series above. |
1793–2019 | Edward F. McQuarrie, “New Bank and Transportation Stock Indexes from
1793 to 1871, with Comparisons Across Region and Sector, and Against Prior
Indexes” (2019),
doi:10.2139/ssrn.3480838. See also “The First 50 Years of the US Stock Market” (ssrn.3209440) and “Stock Market Charts You Never Saw” (ssrn.3050736), which carry his series forward to the present. |
| Cowles Commission stock indexes The standard pre-CRSP US index, with dividends. |
1871–1938 | Yale ICF —
Cowles Data. Alfred Cowles 3rd and Associates, Common-Stock Indexes, 1871–1937, Cowles Commission Monograph No. 3 (Bloomington: Principia Press, 2nd ed. 1939). |
| Sister exchange projects Price data for the other centres, at the ICF. |
1720–1941 | London · St Petersburg · Shanghai |
| Sister news sites The same method applied to the financial press of other centres. |
1850–2008 | London · St Petersburg · Shanghai |