Charleston business directory

The data, whole

Every record behind the directory pages, as files, with what each field means and how accurate the set was measured to be.

Files

Built 2026-09-01 from the same data the pages are built from. CSV for a spreadsheet, JSON for a program; the JSON carries the counts, the street records and the accuracy figures as well.

FileSizeSHA-256 (first 16)
charleston-1852-directory.csv609 KB07ebf8abcaab8ef5
charleston-1852-directory.json2,025 KBfb6120bc09dd5bfa
charleston-1861-register.csv460 KB7e047d2d7371ed22
charleston-1861-register.json1,065 KB28224db2aeb442f0
charleston-1854-1858-advertisers.csv46 KB5a53855edaa9941e
charleston-1854-1858-advertisers.json126 KB27d0d2e0aa29487a
charleston-directory-wikipedia-links.json8 KBb948f4b1c46f65bb

The printed books are in the public domain. This transcription is published by Charleston in the War for any use with attribution to charlestoninthewar.com; the Wikipedia sentences in the links file are CC BY-SA 4.0 as marked.

What is in it

1852. 5,655 entries from Bagget's directory, 5,121 with a place of business, on 158 streets. 1861. 4,248 houses from Ford's census, owner and occupant, 169 of them plainly a business. 1854 and 1858. 194 dated advertisements naming 179 houses. Wikipedia. 15 links, each with the article's own sentence.

How accurate it is

Measured, not assumed. Thirty 1852 entries drawn at random from ten random pages (three per page) and twenty 1861 houses from five random pages (four per page), each read against the page image by eye.

1852 directory. 15 of 30 entries were right in every field; 27 of 30 carried the right address; 3 took the wrong address; 10 had a fault in the name; 2 sat under the wrong surname. Half the sample is right in every field. Nine in ten carry the right address. The commonest fault is a name split wrong — a forename filed as the trade, or a title without its initial — and the sample turned up two people filed under a neighbour's surname where a ditto mark was lost. The systematic faults were then corrected across the whole file (96 forenames, 190 titles) and the doubtful surnames flagged; the figures above describe the file as it stood before those corrections.

1861 register. 19 of 20 houses were on the right street and 18 had the right number; the side of the street was right on 12; 2 were right in every field. The census scan is stained and its machine reading is poor: the street and number are nearly always right, but the side of the street was wrong on eight houses and the owner or occupant text is garbled or wrongly split on twelve. The one wrong street (Pritchard filed under King) was corrected and the side of every house was then re-read from the page headers, after which 18 of the 20 audited houses stood on the right side; the owner and occupant text stands as the machine read it.

Second scan. A second scan of the 1852 book, read independently, agrees with 5,452 of 5,655 lines; where it differs, both readings are in the file.

What to trust. The raw line and the page reference are the ground truth in every file; every other field is a reading of them. A flag on a record means the reading departed from the plain line and says how. Nothing was corrected without the page image or the source line, and nothing was corrected silently.

Fields, 1852

surname / forename
As the directory prints them, cased for reading. A ditto mark in the book means the surname above; the Mc section prints 'Mc' once and dittos after it (flag mc-from-ditto).
trade
The occupation as printed, lower case. Empty where the book gives none (flag no-occupation-in-source) or where it is dittoed from the line above (occupation-dittoed).
address / street / number / kind
The place of business as printed; street is the normalised key that the street pages use; kind is numbered, wharf, corner or named. Never a modern location: Charleston renumbered its streets after 1882.
residence
The second address on the line, where the book gives one.
ward
The ward, where the book gives one instead of a number.
flags
Every departure from a plain reading, semicolon-separated: page-checked, forename-restored, surname-from-firm-line, mc-from-ditto, restored-from-source-line, surname-uncertain (with surname_candidate in the JSON), number-split-by-ocr, street-name-repaired, occupation-too-long, firm, no-address.
scan2_grade / scan2_text
The verdict of a second, independent scan of the same book aligned line by line: agree, close, differ or unaligned; the second scan's text where it differs.
raw
The machine reading of the printed line, untouched. It is the thing to check every other field against.

Fields, 1861

street / number / side / ward
From the page headers of Ford's census, re-read from the page images on 1 September 2026 (the first machine reading filed whole streets under their neighbours). street_as_parsed keeps the first reading where it differed.
owner / occupant
As the machine read the two columns. The scan is stained and these fields are the least reliable in the set; where the columns were re-split the first reading is kept in owner_as_parsed / occupant_as_parsed and note says so.
is_business
True where the occupant is plainly a firm, an office or an institution rather than a person.
leaf
The page of the scanned book the row was read from, so any row can be checked.

Fields, 1854 and 1858

One row per advertisement: the firm as printed, the year, the source and its page, the trade and address as the card gives them, the principals named, the kind of notice (card, editorial, section, elsewhere, uncertain), and printed_as where the OCR damaged the name. A listing proves a year, not a span.

Sources: Directory of the City of Charleston, for the Year 1852; Census of the City of Charleston, South Carolina, for the Year 1861; the 1854 and 1858 volumes as named on the business pages.