Dataset Viewer
Auto-converted to Parquet Duplicate
image
image
id
string
title
string
date
string
language
string
decade
int32
collection
string
data_provider
string
rights
string
page_iiif_url
string
page
int32
text
string
mean_ocr
float64
width
int32
height
int32
alto_unit
string
alto_page_size
list
alto_image_aspect_diff
float64
box_alignment
string
lines
list
words
list
alto_xml
string
stratum
string
stratum_rank
int64
tier
int32
https://www.europeana.eu/item/9200300/BibliographicResource_3000051796867/$3
Wiener Zeitung
1705-02-25
de
1,700
9200300
Österreichische Nationalbibliothek - Austrian National Libra
http://creativecommons.org/publicdomain/mark/1.0/
https://iiif.onb.ac.at/i…ll/0/default.jpg
3
' herein sich wieder begeben / da man dieselbe abermahl gesanglich eingezogen / !i ihr den erimmal-krocesskormiret/undzum andern mahl Mit einem gantzcn Z Schilling das Land auff ewig verwiesen/mit diesem Beysatz / dasern sie die 3 Oesterreichische Landen wieder betretten würde/ mit dem Schwerdt hingerich tet werden sol...
0.522676
1,760
2,456
pixel
[ 1780, 2456 ]
0.01124
ok
[ { "text": "' herein sich wieder begeben / da man dieselbe abermahl gesanglich eingezogen /", "bbox": [ 14.8, 187, 1553.3, 265 ], "mean_wc": 0.4886 }, { "text": "!i ihr den erimmal-krocesskormiret/undzum andern mahl Mit einem gantzcn", "bbox": [ 5.9, 24...
[ { "text": "'", "bbox": [ 14.8, 192, 24.7, 199 ], "wc": 0.6299999952, "line": 0 }, { "text": "herein", "bbox": [ 40.5, 192, 156.2, 250 ], "wc": 0.5083333254, "line": 0 }, { "text": "sich", "bbox": [ 168.1, ...
<?xml version="1.0" encoding="UTF-8" standalone="yes"?> <alto xmlns="http://www.loc.gov/standards/alto/ns-v2#" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.loc.gov/standards/alto/ns-v2# http://www.loc.gov/standards/alto/alto-v2.0.xsd"> ...
9200300|de|1700
1
0
https://www.europeana.eu/item/9200300/BibliographicResource_3000051795724/$7
Wiener Zeitung
1708-07-25
de
1,700
9200300
Österreichische Nationalbibliothek - Austrian National Libra
http://creativecommons.org/publicdomain/mark/1.0/
https://iiif.onb.ac.at/i…ll/0/default.jpg
7
"Änkunfft derer Hoch-und nideren Stands^Persohnen. Den 25. Zulii r/oS. Rothen, Thurn, He« General (...TRUNCATED)
0.51519
1,848
2,536
pixel
[ 1864, 2536 ]
0.00858
ok
[{"text":"Änkunfft derer Hoch-und nideren Stands^Persohnen.","bbox":[192.3,223.0,1420.7,320.0],"mea(...TRUNCATED)
[{"text":"Änkunfft","bbox":[193.3,223.0,401.5,293.0],"wc":0.5762500167,"line":0},{"text":"derer","b(...TRUNCATED)
"<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?>\r\n<alto xmlns=\"http://www.loc.gov/s(...TRUNCATED)
9200300|de|1700
2
0
https://www.europeana.eu/item/9200300/BibliographicResource_3000116312586/$7
Wiener Zeitung
1709-10-02
de
1,700
9200300
Österreichische Nationalbibliothek - Austrian National Libra
http://creativecommons.org/publicdomain/mark/1.0/
https://iiif.onb.ac.at/i…ll/0/default.jpg
7
"i Dtm Philipp Grdkager/ und Agnes feiner Ehew. ihr Zwilling/Gabriel und Maria EvaRosiua. Dem Johann(...TRUNCATED)
0.502916
1,877
2,369
pixel
[ 1877, 2369 ]
0
ok
[{"text":"i","bbox":[1530.0,2137.0,1535.0,2158.0],"mean_wc":0.22},{"text":"Dtm Philipp Grdkager/ und(...TRUNCATED)
[{"text":"i","bbox":[1530.0,2137.0,1535.0,2158.0],"wc":0.2199999988,"line":0},{"text":"Dtm","bbox":[(...TRUNCATED)
"<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?>\r\n<alto xmlns=\"http://www.loc.gov/s(...TRUNCATED)
9200300|de|1700
9
1
https://www.europeana.eu/item/9200300/BibliographicResource_3000051870221/$7
Wiener Zeitung
1705-11-25
de
1,700
9200300
Österreichische Nationalbibliothek - Austrian National Libra
http://creativecommons.org/publicdomain/mark/1.0/
https://iiif.onb.ac.at/i…ll/0/default.jpg
7
"ferl. Currier / kombtvon Münches /log. im Posi-Hauß. Den'26. Dito. Stuben-Thor. Her! Frantz Reymu(...TRUNCATED)
0.542771
1,752
2,336
pixel
[ 1752, 2288 ]
0.02055
unverified
[{"text":"ferl. Currier / kombtvon Münches /log.","bbox":[198.0,74.5,889.0,123.5],"mean_wc":0.6327}(...TRUNCATED)
[{"text":"ferl.","bbox":[198.0,78.6,265.0,120.5],"wc":0.6240000129,"line":0},{"text":"Currier","bbox(...TRUNCATED)
"<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?>\r\n<alto xmlns=\"http://www.loc.gov/s(...TRUNCATED)
9200300|de|1700
10
1
https://www.europeana.eu/item/9200300/BibliographicResource_3000051825119/$8
Wiener Zeitung
1706-03-31
de
1,700
9200300
Österreichische Nationalbibliothek - Austrian National Libra
http://creativecommons.org/publicdomain/mark/1.0/
https://iiif.onb.ac.at/i…ll/0/default.jpg
8
"werde; Undseye gewiß/baß der KKlorä ?eterborou» 6. Fratttz5stsche ttillons geschlagen/ hwgegen (...TRUNCATED)
0.484776
1,860
2,500
pixel
[ 1860, 2500 ]
0
ok
[{"text":"werde; Undseye gewiß/baß der KKlorä ?eterborou» 6. Fratttz5stsche","bbox":[219.0,187.0(...TRUNCATED)
[{"text":"werde;","bbox":[219.0,195.0,371.0,243.0],"wc":0.3350000083,"line":0},{"text":"Undseye","bb(...TRUNCATED)
"<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?>\r\n<alto xmlns=\"http://www.loc.gov/s(...TRUNCATED)
9200300|de|1700
3
1
https://www.europeana.eu/item/9200300/BibliographicResource_3000116310668/$3
Wiener Zeitung
1709-12-28
de
1,700
9200300
Österreichische Nationalbibliothek - Austrian National Libra
http://creativecommons.org/publicdomain/mark/1.0/
https://iiif.onb.ac.at/i…ll/0/default.jpg
3
"Auß Poh/ett/von dem sa December. Daß die Stadt Riga/ wke voll Wildav.-r/autt/ mit Feuer starck be(...TRUNCATED)
0.439715
1,845
2,363
pixel
[ 1845, 2363 ]
0
ok
[{"text":"Auß Poh/ett/von dem sa December. Daß die Stadt Riga/ wke","bbox":[319.0,99.0,1699.0,200.(...TRUNCATED)
[{"text":"Auß","bbox":[319.0,103.0,416.0,155.0],"wc":0.453333348,"line":0},{"text":"Poh/ett/von","b(...TRUNCATED)
"<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?>\r\n<alto xmlns=\"http://www.loc.gov/s(...TRUNCATED)
9200300|de|1700
11
1
https://www.europeana.eu/item/9200300/BibliographicResource_3000051867403/$8
Wiener Zeitung
1706-04-07
de
1,700
9200300
Österreichische Nationalbibliothek - Austrian National Libra
http://creativecommons.org/publicdomain/mark/1.0/
https://iiif.onb.ac.at/i…ll/0/default.jpg
8
"Niederlanden echieltm / marschirten an dem Mosel- und Saarsirohm bald c hin / bald her / und käme (...TRUNCATED)
0.530367
1,860
2,488
pixel
[ 1860, 2488 ]
0
ok
[{"text":"Niederlanden echieltm / marschirten an dem Mosel- und Saarsirohm bald c","bbox":[302.0,188(...TRUNCATED)
[{"text":"Niederlanden","bbox":[302.0,188.0,584.0,241.0],"wc":0.3574999869,"line":0},{"text":"echiel(...TRUNCATED)
"<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?>\r\n<alto xmlns=\"http://www.loc.gov/s(...TRUNCATED)
9200300|de|1700
12
1
https://www.europeana.eu/item/9200300/BibliographicResource_3000051780711/$7
Wiener Zeitung
1707-04-23
de
1,700
9200300
Österreichische Nationalbibliothek - Austrian National Libra
http://creativecommons.org/publicdomain/mark/1.0/
https://iiif.onb.ac.at/i…ll/0/default.jpg
7
"DtS»5.dlt». sterischen Regiment / kombt von Eärnk«er»TH«r. Heri Obrist Wachtmei« bürg/log. (...TRUNCATED)
0.487461
1,883
2,443
pixel
[ 1883, 2443 ]
0
ok
[{"text":"DtS»5.dlt». sterischen Regiment / kombt von","bbox":[367.0,243.0,1492.0,303.0],"mean_wc"(...TRUNCATED)
[{"text":"DtS»5.dlt».","bbox":[367.0,243.0,597.0,288.0],"wc":0.5790908933,"line":0},{"text":"steri(...TRUNCATED)
"<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?>\r\n<alto xmlns=\"http://www.loc.gov/s(...TRUNCATED)
9200300|de|1700
14
1
https://www.europeana.eu/item/9200300/BibliographicResource_3000051814448/$6
Wiener Zeitung
1709-04-10
de
1,700
9200300
Österreichische Nationalbibliothek - Austrian National Libra
http://creativecommons.org/publicdomain/mark/1.0/
https://iiif.onb.ac.at/i…ll/0/default.jpg
6
"gcnwa'rtigen Reichs-Tag gelangenNicht weniger Ihro Kayserl. Majestät sothanes Werck/durchein Kayse(...TRUNCATED)
0.535429
1,800
2,428
pixel
[ 1800, 2428 ]
0
ok
[{"text":"gcnwa'rtigen Reichs-Tag gelangenNicht weniger Ihro Kayserl. Majestät","bbox":[174.0,111.0(...TRUNCATED)
[{"text":"gcnwa'rtigen","bbox":[174.0,132.0,425.0,189.0],"wc":0.3700000048,"line":0},{"text":"Reichs(...TRUNCATED)
"<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?>\r\n<alto xmlns=\"http://www.loc.gov/s(...TRUNCATED)
9200300|de|1700
13
1
https://www.europeana.eu/item/9200300/BibliographicResource_3000051830238/$10
Wiener Zeitung
1706-08-21
de
1,700
9200300
Österreichische Nationalbibliothek - Austrian National Libra
http://creativecommons.org/publicdomain/mark/1.0/
https://iiif.onb.ac.at/i…ll/0/default.jpg
10
"Ver Elisabeth Bürgerin / Wittib / im ersteche« / vnd von bannen ins Bartz rothenEreutzamLerchen-F(...TRUNCATED)
0.530182
1,879
2,496
pixel
[ 1879, 2496 ]
0
ok
[{"text":"Ver Elisabeth Bürgerin / Wittib / im ersteche« / vnd von bannen ins Bartz","bbox":[270.0(...TRUNCATED)
[{"text":"Ver","bbox":[270.0,183.0,343.0,219.0],"wc":0.2933333218,"line":0},{"text":"Elisabeth","bbo(...TRUNCATED)
"<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?>\r\n<alto xmlns=\"http://www.loc.gov/s(...TRUNCATED)
9200300|de|1700
4
1
End of preview. Expand in Data Studio

Europeana Newspapers: Page Images with OCR Layout

98,877 newspaper page images, 1700s–1940s, in 10 languages (plus pages marked multi-language or unidentified), each paired with the OCR produced when the page was digitised: full text, the ALTO XML, and every line and word box already converted to image pixels with the OCR engine's per-word confidence. The pages come from eight European libraries via Europeana Newspapers and are a sample of the 5.9 million pages in biglam/europeana_newspapers, joined on id.

Status: proof of concept. This is a first test release to see whether page images plus the original OCR layout are useful as a Hub dataset. The sample is designed to grow (see Sampling); the structure may still change.

What it looks like

All three images are from one page: Hamburger Nachrichten, 22 January 1840, page 7 (Hamburg State Library).

OCR line boxes drawn on the page image lines[].bbox drawn on image: the OCR layout, already in image pixels.

Word boxes coloured by OCR confidence words[].bbox coloured by words[].wc, from red (low confidence) to green (high).

One line image with its OCR text One line cut from the page with its OCR text. Even at a mean confidence of 0.79 the OCR reads "Prineipals. Rsflecttrende" for "Principals. Reflectirende": treat the text as silver, not gold.

What's in a row

Column Description
image Full-size page image as served by the library's IIIF server (not re-encoded)
id Europeana page id; joins to biglam/europeana_newspapers (data and alto configs)
title, date, language, decade, collection, data_provider Newspaper title, issue date, language (from the source dataset), Europeana collection id, holding library
text Page text as extracted in biglam/europeana_newspapers
mean_ocr Mean OCR word confidence for the page (from the source dataset)
lines One entry per ALTO TextLine: text, bbox [x0, y0, x1, y1] in image pixels, mean_wc
words One entry per ALTO String: text, bbox in image pixels, wc (OCR confidence 0–1), line index
alto_xml The raw ALTO v2 XML for the page
width, height Image size in pixels
alto_unit, alto_page_size ALTO measurement unit (pixel or mm10) and page size in ALTO units
box_alignment, alto_image_aspect_diff Whether the boxes can be trusted to land on the image; see below
page_iiif_url IIIF URL of the page image
rights Rights statement from the Europeana record (all Public Domain Mark 1.0)
stratum, stratum_rank, tier Sampling bookkeeping (see Sampling)
from datasets import load_dataset

# stream: rows carry full-size scans (~3 MB each)
ds = load_dataset("biglam/europeana_newspapers_images", split="train", streaming=True)
row = next(iter(ds))
row["image"].crop(row["lines"][0]["bbox"])  # first OCR line, cut from the page

Pick pages before downloading. The metadata config is a 7 MB file with one row per page and no images, text or ALTO (id, title, date, language, collection, mean_ocr, mean_wc, box_alignment, image size, line and word counts, page_iiif_url, and data_file, the parquet shard the page is in). Filter it first, then read images only from the shards you need:

import duckdb
from datasets import Dataset, Image

repo = "hf://datasets/biglam/europeana_newspapers_images"
# 1. choose pages from the small metadata file
picked = duckdb.sql(f"""
    select id, data_file from '{repo}/metadata/metadata.parquet'
    where "language" = 'el' and box_alignment = 'ok'
    order by mean_wc limit 5""").df()
# 2. read those pages from their shards (each page costs about one 60 MB row group)
table = duckdb.execute(
    "select * from read_parquet(?) where id in (select unnest(?))",
    [[f"{repo}/{f}" for f in picked.data_file.unique()], picked.id.tolist()],
).to_arrow_table()
ds = Dataset(table).cast_column("image", Image())  # a normal datasets.Dataset

This holds the selected pages in memory (roughly 3–4 MB per page). For more than about a thousand pages, write them to a local parquet file instead and load that; load_dataset memory-maps it rather than holding it in RAM:

from datasets import load_dataset

duckdb.execute(
    "copy (select * from read_parquet(?) where id in (select unnest(?))) to 'subset.parquet'",
    [[f"{repo}/{f}" for f in picked.data_file.unique()], picked.id.tolist()],
)
# COPY drops the image feature type, so cast it back
ds = load_dataset("parquet", data_files="subset.parquet", split="train").cast_column("image", Image())

Possible uses

  • OCR training and evaluation. Crop lines[].bbox from image and pair with lines[].text for line-level OCR data in 10 languages and several scripts (Fraktur, Cyrillic, Greek). The targets are the original OCR, so they are silver, not gold: good for pre-training or for finding hard pages, not as a benchmark truth.
  • OCR quality estimation. words[].wc and lines[].mean_wc give the original engine's confidence at word level, so bad regions of a page can be found and re-OCR'd, or left out, rather than scoring the page as a whole.
  • Re-OCR comparisons. Run a modern OCR or vision-language model on the same pages and compare against the historical OCR, per language, decade or collection.
  • Layout and reading order. Line and word boxes plus the block structure in alto_xml give page layout for multi-column historical newspapers.

Things to know before using it

The OCR is historical and uneven. It was produced by the libraries when the pages were digitised (the ALTO names ABBYY FineReader Engine for 65,636 pages and the CCS docWorks workflow for 33,241) and was taken from Europeana's 2019 full-text dumps. Quality varies by collection, typeface and paper. Some Serbian Cyrillic pages were OCR'd as Latin characters, which leaves unreadable text with plausible-looking boxes.

OCR confidence is not comparable across collections. The same print quality gets very different wc values in different collections (the engine settings differed), so compare confidence within a collection, not across them.

Image resolution depends on the library. "Full size" is whatever each IIIF server provides:

Collection Library Pages Median image width (px) 10th–90th percentile
9200356 National Library of Estonia 22,541 4,000 2,460–6,127
9200300 Austrian National Library 16,278 2,040 1,527–3,206
9200339 University of Belgrade 13,909 1,393 1,019–1,834
9200338 Hamburg State Library 13,343 4,046 2,629–5,112
9200301 National Library of Finland 11,277 2,500 1,856–2,812
9200355 Berlin State Library 8,343 3,702 2,536–4,733
9200396 National Library of Luxembourg 6,776 1,256 1,256–1,256
9200357 National Library of Poland 6,410 2,302 1,621–4,031

Check box_alignment before cropping. Boxes are ALTO coordinates scaled by image size / ALTO page size. That is right when the image and the ALTO describe the same page frame. For 637 pages (610 of them Austrian National Library scans, some of which include the facing page or wide margins) the image and ALTO aspect ratios differ by more than 2%, and the boxes drift off the text; these are marked box_alignment = "unverified". In a visual review of 40 pages, the boxes on the other pages landed on the text.

Some pages were left out on purpose. Pages from collection 9200357 (National Library of Poland) that are in Izraelita or dated 1939 are excluded: for these, the OCR often belongs to a neighbouring page of the issue rather than the image. Pages whose images are no longer served (1,123 pages, about 1% of those tried, almost all from the Austrian National Library) are missing too, which is why there are 98,877 rows rather than 100,000.

Sampling

The sample is not representative of the source dataset, by design. Pages were allocated by language first (share proportional to the square root of each language's page count, with a floor so small languages are included), then across collection × decade within each language, with at most 5% of pages from any one title and, where possible, one page per issue. Pages are ranked deterministically within each stratum, so a larger sample is a strict superset of this one (tier 0 is the original 1,000-page pilot).

Language Pages
German (de) 34,052
Estonian (et) 12,296
Serbian (sr) 11,416
French (fr) 6,776
Polish (pl) 6,410
Finnish (fi) 6,241
multi-language 5,762
Swedish (sv) 5,036
Russian (ru) 4,198
Greek (el) 3,715
no language found 2,493
Croatian (hr) 482

Only pages whose page image could be matched reliably were eligible. The source dataset's item_iiif_url points to each issue's first page; the page-level image links here come from Europeana's 2019 metadata dump (edm:isShownBy + edm:hasView), checked against the live IIIF manifests and the images themselves.

How it was made

Built on Hugging Face Jobs: selection and text/ALTO extraction with DuckDB over biglam/europeana_newspapers, then images fetched from the libraries' IIIF servers at a rate each server was comfortable with (about 3–4 images/s for Europeana's server), staged in an HF bucket, and assembled into parquet. The fetcher identified itself with a User-Agent and backed off when a server slowed down.

Licence

Europeana's records mark all these items Public Domain Mark 1.0, as asserted by the holding libraries (2019 metadata). Some material from the 1910s–1940s may still be in copyright in some jurisdictions despite that mark; check before commercial reuse.

Credit

All newspapers, images and OCR were created by the Austrian National Library, National Library of Finland, Hamburg State Library, University of Belgrade, Berlin State Library, National Library of Estonia, National Library of Poland and National Library of Luxembourg, and aggregated by Europeana Newspapers. Sampled and repackaged by Daniel van Strien (Machine Learning Librarian, Hugging Face).

Downloads last month
487