Characters per Word: Moltoyo Text Density Study

MTD-v1-2026 · Data version dated 26 September 2026 · Page first published 5 October 2026

We measured 250 English text samples of 500 words each: 125,000 words across five text types. In this corpus, the mean was 6.33 characters per word including spaces, ranging from 5.53 for classic fiction to 6.76 for technical and instructional writing.

The results describe this dataset rather than English as a whole. You can inspect text with the Character Counter and Word Counter, or use the practical conversion guide How Many Characters Is…?

Average characters per word by text type

Across all 250 samples the mean was 6.33 characters per word, spaces included. Predicting 3,000 characters for every 500-word sample (the six-character rule) gave a mean absolute percentage error (MAPE) of 9.03%. Values below are rounded from the downloadable measurements.

Mean characters per word (spaces included) and MAPE of the six-character rule
Text typeSamplesMean chars/wordMAPE vs 6
News506.306.41%
General informational506.336.75%
Academic506.7310.67%
Technical & instructional506.7612.41%
Classic fiction505.538.92%
Overall2506.339.03%

Source metadata review — 5 October 2026

  • The downloads contain measurements and source links only. No original prose is included.
  • 91 Europe PMC articles have been confirmed at article level: 89 CC BY 4.0 and 2 CC0.
  • Eight of the 15 Wikinews samples have incorrect recorded publication dates. These include seven sources first published in 2023–2024, before the study's cutover date. Those seven should be recorded as CC BY 2.5 under Wikinews:Copyright, not the recorded CC BY 4.0. The samples and measurements remain in v1; details are in metadata-errata.json.
  • Other attribution and jurisdiction questions remain under review; the measurements are unchanged.

How we measured characters per word

Corpus

250 samples, 50 in each of five categories: news, general informational, academic, technical and instructional, and classic fiction. Each sample is 500 words. Every source document had at least 750 eligible words after cleaning; the smallest accepted document had 757.

Normalization (MTD-NORM-1)

Non-prose material such as references, navigation, figures, tables and abstracts was removed. All runs of whitespace were collapsed to a single space before sampling. Unicode characters were preserved.

Selection

Within each eligible quota pool, documents were ranked by SHA-256(seed|version|rank|source_document_id), with no more than two fiction works per author. The seed is MOLTOYO-TEXT-DENSITY-V1-2026 and the version is exactly MTD-v1-2026.

Sampling

Each sample is 500 contiguous whitespace-separated tokens. With N eligible cleaned tokens, the start skips the leading 5% where possible, then adds a hash-derived offset that can land anywhere from that floor up to the last possible start (hence the +1):

S     = N - 500
floor = min(floor(N * 0.05), S)
U     = S - floor
h     = BigInt('0x' + sha256(seed + '|' + version + '|start|' + source_document_id).slice(0, 16))
start = U <= 0 ? floor : floor + Number(h % BigInt(U + 1))
end   = start + 499   // inclusive

Locators in the data are zero-based token indices, and the end locator is inclusive.

Measurement

A word is a token after trimming and splitting on /\s+/. Characters are Unicode code points (counted with Array.from), spaces included. Characters per word is characters ÷ 500. The six-character rule predicts 3,000 characters per sample; absolute percentage error is 100 × |actual − 3,000| ÷ actual, and MAPE is the average of those errors.

These are the frozen 2026 study definitions. Moltoyo's live counting tools were updated in October 2026 to count user-perceived characters and words containing letters or numbers, so their results can differ from the study's definitions.

Sources and revisions

Final source mix (samples per source)
CategorySources
News15 Wikinews, 35 Global Voices
General informational32 Wikipedia, 17 GOV.UK, 1 European Commission
Academic50 Europe PMC (10 each: biomedical & clinical, life sciences, physical & earth sciences, engineering & technology, social & behavioral sciences)
Technical & instructional41 Europe PMC-based CC BY technical publications, 9 GOV.UK
Classic fiction50 Project Gutenberg, no more than 2 works per author

Documented changes to the planned source mix

  1. A second news source was added: 35 Global Voices samples alongside Wikinews.
  2. SciDev.Net was excluded because its material could not be verified.
  3. NIST material was unavailable at the required threshold.
  4. The European Commission quota fell from 13 to 1; the rest was reassigned to Wikipedia and GOV.UK.
  5. The GOV.UK technical quota fell from 22 to 9; the rest was reassigned to 41 CC BY technical publications.

Not every requirement of the original protocol was met; see the metadata review above. The full list of 250 source links is in sources.json.

Download the study data

Downloads contain measurements and metadata only. Source works retain their own rights; no licence is granted here over any source content.

  • measurements.csv

    CSV, UTF-8

    250 rows of per-sample measurements

  • measurements.json

    JSON, UTF-8

    The same 250 rows with numeric types

  • sources.json

    JSON, UTF-8

    250 source links: sample ID, source and URL

  • methodology.json

    JSON, UTF-8

    Frozen definitions, final source mix, revisions and the dated metadata review

  • metadata-errata.json

    JSON, UTF-8

    8 source date and licence corrections from the 5 October 2026 review. Metadata only, no prose; the original study records are unchanged

Field dictionary

Columns in measurements.csv and measurements.json.

FieldMeaning
public_sample_idStable sample identifier: MTD-v1-2026/ plus a source prefix and source document ID.
categoryOne of news, general_informational, academic, technical_instructional, classic_fiction.
subdisciplineAcademic samples only (empty otherwise).
moltoyo_wordsWords under the study definition. Always 500.
characters_with_spacesUnicode code points in the sample, spaces included.
characters_without_whitespaceCode points excluding whitespace.
characters_per_moltoyo_word_with_spacescharacters_with_spaces ÷ 500.
absolute_percentage_error_vs_6_rule100 × |actual − 3,000| ÷ actual, as a percentage.
sample_sha256SHA-256 of the normalized sample. It identifies the sample; it is not proof of rights or licence.
sample_start_locatorZero-based index of the first token in the normalized source document.
sample_end_locatorZero-based index of the last token (inclusive).

Limitations

  • The results describe this corpus, not English as a whole.
  • The 750-word minimum favours long-form documents.
  • Academic and technical samples lean heavily on biomedical and scientific publications from Europe PMC.
  • Each category has 50 samples; no confidence intervals are reported.
  • The study's word and character definitions differ from those used by Moltoyo's live tools.
  • Source metadata review is ongoing, as described above.

How to cite this study

Data version MTD-v1-2026 is dated 26 September 2026. This page was first published on 5 October 2026.

Moltoyo. (2026). Moltoyo Text Density Study, MTD-v1-2026. https://moltoyo.com/research/text-density-study