4K & HD English Speech Video + Audio Dataset for AI Training
Access a video and audio dataset containing 1,126 horizontal clips with recorded English speech, delivered in both 4K/UHD and HD. The collection provides approximately 15.4 hours of synchronized audiovisual content for AI training, fine-tuning, evaluation, enrichment, and commercial model-development workflows.
The dataset combines visible human-centered video with spoken English audio, making it suitable for multimodal systems that need to interpret relationships between speech, speakers, facial and body cues, and visual context. Structured metadata supports retrieval, filtering, organization, and integration into AI development pipelines. Synchronized speech and video of this kind sits within our wider range of multimodal AI datasets.
Dataset Preview
Representative previews showing English-language speech captured on video with synchronized recorded audio across varied human-centered scenes.
The examples illustrate visible speakers, speaking contexts, and the audio-visual alignment available for speech, lip-motion, and multimodal learning workflows.
Key Highlights
- • Video clips are paired with recorded English speech audio for synchronized audiovisual learning
- • Human-centered video suitable for audiovisual model development
- • Structured metadata for filtering, retrieval, organization, and AI workflow support
- • No subtitles or scripts supplied; English audio can be transcribed on request
Metadata Fields
Technical Specifications
Dataset Type
Video + Audio Dataset: English Speech Horizontal
Content Type
Video clips paired with recorded English speech audio
Clip Count
1,126 clips
Total Duration
15.4 hours
Total Duration in Minutes
Approximately 926.2 minutes
Average Clip Length
Approximately 49.4 seconds
Resolution
UHD: 3840 × 2160
DCI 4K: 4096 × 2160 where available
Full HD: 1920 × 1080
Formats
4K/UHD + HD
Codec
H.264
4K/UHD Bitrate
Approximately 87–93 Mbps
HD Bitrate
Approximately 22 Mbps
Container
MP4
Orientation
Horizontal
Aspect Ratio Mixed horizontal formats
16:9 UHD/HD; 256:135 DCI 4K where available
Spoken Language
English
Audio Content
Recorded speech audio paired with video
Subtitles / Scripts
N/A, no subtitles or scripts supplied
Transcription
English audio transcribable on request
Licensing, Documentation, and Delivery
Wavebreak Media can provide applicable licensing, provenance, and model or property release information for the audiovisual content, with dataset delivery organized around agreed video, audio, metadata, transcription, and technical requirements. See our Dataset Licensing & Compliance and Dataset Delivery & Security pages for details on rights documentation, packaging, and secure transfer.
Need the Full Dataset?
Request access to the full 4K and HD English speech video and audio dataset or discuss a custom audiovisual dataset requirement with Wavebreak Media.
Related Datasets
4K & HD Object & Foley Sounds Video + Audio Dataset for AI Training
Access 25,657 video clips with recorded object and foley sounds in 4K and HD for multimodal AI training, audiovisual understanding, sound recognition, and evaluation.
View Dataset VIDEO + AUDIO DATASETEveryday Object Sounds Dataset
Access 25,000 training-ready 4K video clips with synchronized object and material sounds for audio-visual AI, foley generation, sound synthesis, and multimodal learning.
View Dataset 2-MINUTE END-TO-END TASK CAPTURELong-Form Human Procedural Video Dataset
Access 579+ long-form 4K procedural videos capturing complete human tasks for sequential action recognition, procedural reasoning, robotics, and embodied AI.
View Dataset
