Characters per Word: Moltoyo Text Density Study
MTD-v1-2026 · Data version dated 26 September 2026 · Page first published 5 October 2026
We measured 250 English text samples of 500 words each: 125,000 words across five text types. In this corpus, the mean was 6.33 characters per word including spaces, ranging from 5.53 for classic fiction to 6.76 for technical and instructional writing.
The results describe this dataset rather than English as a whole. You can inspect text with the Character Counter and Word Counter, or use the practical conversion guide How Many Characters Is…?
Average characters per word by text type
Across all 250 samples the mean was 6.33 characters per word, spaces included. Predicting 3,000 characters for every 500-word sample (the six-character rule) gave a mean absolute percentage error (MAPE) of 9.03%. Values below are rounded from the downloadable measurements.
| Text type | Samples | Mean chars/word | MAPE vs 6 |
|---|---|---|---|
| News | 50 | 6.30 | 6.41% |
| General informational | 50 | 6.33 | 6.75% |
| Academic | 50 | 6.73 | 10.67% |
| Technical & instructional | 50 | 6.76 | 12.41% |
| Classic fiction | 50 | 5.53 | 8.92% |
| Overall | 250 | 6.33 | 9.03% |
Source metadata review — 5 October 2026
- The downloads contain measurements and source links only. No original prose is included.
- 91 Europe PMC articles have been confirmed at article level: 89 CC BY 4.0 and 2 CC0.
- Eight of the 15 Wikinews samples have incorrect recorded publication dates. These include seven sources first published in 2023–2024, before the study's cutover date. Those seven should be recorded as CC BY 2.5 under Wikinews:Copyright, not the recorded CC BY 4.0. The samples and measurements remain in v1; details are in metadata-errata.json.
- Other attribution and jurisdiction questions remain under review; the measurements are unchanged.
How we measured characters per word
Corpus
250 samples, 50 in each of five categories: news, general informational, academic, technical and instructional, and classic fiction. Each sample is 500 words. Every source document had at least 750 eligible words after cleaning; the smallest accepted document had 757.
Normalization (MTD-NORM-1)
Non-prose material such as references, navigation, figures, tables and abstracts was removed. All runs of whitespace were collapsed to a single space before sampling. Unicode characters were preserved.
Selection
Within each eligible quota pool, documents were ranked by SHA-256(seed|version|rank|source_document_id), with no more than two fiction works per author. The seed is MOLTOYO-TEXT-DENSITY-V1-2026 and the version is exactly MTD-v1-2026.
Sampling
Each sample is 500 contiguous whitespace-separated tokens. With N eligible cleaned tokens, the start skips the leading 5% where possible, then adds a hash-derived offset that can land anywhere from that floor up to the last possible start (hence the +1):
S = N - 500
floor = min(floor(N * 0.05), S)
U = S - floor
h = BigInt('0x' + sha256(seed + '|' + version + '|start|' + source_document_id).slice(0, 16))
start = U <= 0 ? floor : floor + Number(h % BigInt(U + 1))
end = start + 499 // inclusiveLocators in the data are zero-based token indices, and the end locator is inclusive.
Measurement
A word is a token after trimming and splitting on /\s+/. Characters are Unicode code points (counted with Array.from), spaces included. Characters per word is characters ÷ 500. The six-character rule predicts 3,000 characters per sample; absolute percentage error is 100 × |actual − 3,000| ÷ actual, and MAPE is the average of those errors.
These are the frozen 2026 study definitions. Moltoyo's live counting tools were updated in October 2026 to count user-perceived characters and words containing letters or numbers, so their results can differ from the study's definitions.
Sources and revisions
| Category | Sources |
|---|---|
| News | 15 Wikinews, 35 Global Voices |
| General informational | 32 Wikipedia, 17 GOV.UK, 1 European Commission |
| Academic | 50 Europe PMC (10 each: biomedical & clinical, life sciences, physical & earth sciences, engineering & technology, social & behavioral sciences) |
| Technical & instructional | 41 Europe PMC-based CC BY technical publications, 9 GOV.UK |
| Classic fiction | 50 Project Gutenberg, no more than 2 works per author |
Documented changes to the planned source mix
- A second news source was added: 35 Global Voices samples alongside Wikinews.
- SciDev.Net was excluded because its material could not be verified.
- NIST material was unavailable at the required threshold.
- The European Commission quota fell from 13 to 1; the rest was reassigned to Wikipedia and GOV.UK.
- The GOV.UK technical quota fell from 22 to 9; the rest was reassigned to 41 CC BY technical publications.
Not every requirement of the original protocol was met; see the metadata review above. The full list of 250 source links is in sources.json.
Download the study data
Downloads contain measurements and metadata only. Source works retain their own rights; no licence is granted here over any source content.
- measurements.csv
CSV, UTF-8
250 rows of per-sample measurements
- measurements.json
JSON, UTF-8
The same 250 rows with numeric types
- sources.json
JSON, UTF-8
250 source links: sample ID, source and URL
- methodology.json
JSON, UTF-8
Frozen definitions, final source mix, revisions and the dated metadata review
- metadata-errata.json
JSON, UTF-8
8 source date and licence corrections from the 5 October 2026 review. Metadata only, no prose; the original study records are unchanged
Field dictionary
Columns in measurements.csv and measurements.json.
| Field | Meaning |
|---|---|
| public_sample_id | Stable sample identifier: MTD-v1-2026/ plus a source prefix and source document ID. |
| category | One of news, general_informational, academic, technical_instructional, classic_fiction. |
| subdiscipline | Academic samples only (empty otherwise). |
| moltoyo_words | Words under the study definition. Always 500. |
| characters_with_spaces | Unicode code points in the sample, spaces included. |
| characters_without_whitespace | Code points excluding whitespace. |
| characters_per_moltoyo_word_with_spaces | characters_with_spaces ÷ 500. |
| absolute_percentage_error_vs_6_rule | 100 × |actual − 3,000| ÷ actual, as a percentage. |
| sample_sha256 | SHA-256 of the normalized sample. It identifies the sample; it is not proof of rights or licence. |
| sample_start_locator | Zero-based index of the first token in the normalized source document. |
| sample_end_locator | Zero-based index of the last token (inclusive). |
Limitations
- The results describe this corpus, not English as a whole.
- The 750-word minimum favours long-form documents.
- Academic and technical samples lean heavily on biomedical and scientific publications from Europe PMC.
- Each category has 50 samples; no confidence intervals are reported.
- The study's word and character definitions differ from those used by Moltoyo's live tools.
- Source metadata review is ongoing, as described above.
How to cite this study
Data version MTD-v1-2026 is dated 26 September 2026. This page was first published on 5 October 2026.
Moltoyo. (2026). Moltoyo Text Density Study, MTD-v1-2026. https://moltoyo.com/research/text-density-study