Jan Švec | honzas.cz
Menu

# Digital Humanities

5 items

All tags

From UWebASR to calibrated accuracy estimates for oral histories

July 3, 2026

Automatic speech recognition is most useful when users can also estimate how much trust to place in its output. For long witness testimonies, including Holocaust survivor testimonies, this matters because researchers often work with sensitive, historically valuable material. Agentic AI helped us turn a proof-of-concept workflow into a deployable production feature in roughly two weeks.

Portrait preview of the GitHub repository with the UWebASR confidence calibration method

LLM-Based Metadata Extraction from NATO Scanned Documents

June 2, 2026

Historical archives contain valuable evidence, but scanned documents are difficult to search when their metadata is incomplete or inconsistent. At the C4DHI Anniversary Workshop, I presented a workflow that uses large language models to extract structured metadata from scanned NATO archival documents. The talk focused on noisy OCR, multilingual records and the need to preserve evidence for human review.

Portrait preview of the C4DHI workshop programme

Agentic AI for Digital Humanities

26 April - 12 May 2026

Complex research tasks do not fit into a single prompt: they need tools, intermediate checks and an inspectable sequence of steps. These workshop materials introduce agentic AI as a workflow for digital humanities and archival research. They connect an Oxford research stay, CLARIN collaboration and practical work with NATO archival documents.

Portrait preview of the Agentic AI Introduction slides

When speech recognition leaves the lab: UWebASR in practice

July 25, 2025

On 25 July 2025, CLARIN published the impact story Innovative Tool Transforms the Use of Voice Technology. The article describes how UWebASR and related speech and NLP tools from the Czech CLARIN node support media, hackathons, security-oriented processing and academic use. For me, the important point is practical: speech-recognition infrastructure becomes useful when people can rely on it outside a lab demo.

Wide preview with Martin Bulín, Pavel Ircing and Jan Švec from the CLARIN impact story about UWebASR

Asking Questions: finding answers in long oral-history recordings

July 13, 2023

Long oral-history recordings are rich historical sources, but their length and structure make them difficult to search and reuse. On 13 July 2023, CLARIN published the impact story State-of-the-Art Speech Recognition for Understanding Oral Histories, presenting the Semantic Search / Asking Questions framework developed around speech recognition, generated questions and semantic search. The practical goal is to help researchers and visitors ask a question, find the relevant passage and play the answer directly from the original testimony.

Wide preview of the Semantic Search / Asking Questions interface used in the CLARIN impact story