VIDEO + AUDIO DATASET

4K & HD English Speech Video + Audio Dataset for AI Training

Access a video and audio dataset containing 1,126 horizontal clips with recorded English speech, delivered in both 4K/UHD and HD. The collection provides approximately 15.4 hours of synchronized audiovisual content for AI training, fine-tuning, evaluation, enrichment, and commercial model-development workflows.

The dataset combines visible human-centered video with spoken English audio, making it suitable for multimodal systems that need to interpret relationships between speech, speakers, facial and body cues, and visual context. Structured metadata supports retrieval, filtering, organization, and integration into AI development pipelines. Synchronized speech and video of this kind sits within our wider range of multimodal AI datasets.

1,126
Clips
15.4
Total Hours
49.4s
Average Clip Length
4K + HD
Formats
Horizontal
Orientation
English
Spoken Language

Dataset Preview

Representative previews showing English-language speech captured on video with synchronized recorded audio across varied human-centered scenes.

The examples illustrate visible speakers, speaking contexts, and the audio-visual alignment available for speech, lip-motion, and multimodal learning workflows.

Key Highlights

  • Video clips are paired with recorded English speech audio for synchronized audiovisual learning
  • Human-centered video suitable for audiovisual model development
  • Structured metadata for filtering, retrieval, organization, and AI workflow support
  • No subtitles or scripts supplied; English audio can be transcribed on request

Metadata Fields

title Descriptive title of the video clip
description Natural-language description of the visible video content
keywords Keyword tags describing people, activities, environments, subjects, and concepts in the clip
resolution Video resolution, including 4K/UHD or HD
duration Length of the individual audiovisual clip
fps Frames per second
orientation Video orientation; horizontal in this dataset
shoot_date Date associated with the original recording where available
model_release Flag indicating available model-release status where applicable
property_release Flag indicating available property-release status where applicable

Technical Specifications

Dataset Type

Video + Audio Dataset: English Speech Horizontal

Content Type

Video clips paired with recorded English speech audio

Clip Count

1,126 clips

Total Duration

15.4 hours

Total Duration in Minutes

Approximately 926.2 minutes

Average Clip Length

Approximately 49.4 seconds

Resolution

UHD: 3840 × 2160
DCI 4K: 4096 × 2160 where available
Full HD: 1920 × 1080

Formats

4K/UHD + HD

Codec

H.264

4K/UHD Bitrate

Approximately 87–93 Mbps

HD Bitrate

Approximately 22 Mbps

Container

MP4

Orientation

Horizontal

Aspect Ratio Mixed horizontal formats

16:9 UHD/HD; 256:135 DCI 4K where available

Spoken Language

English

Audio Content

Recorded speech audio paired with video

Subtitles / Scripts

N/A, no subtitles or scripts supplied

Transcription

English audio transcribable on request

Licensing, Documentation, and Delivery

Wavebreak Media can provide applicable licensing, provenance, and model or property release information for the audiovisual content, with dataset delivery organized around agreed video, audio, metadata, transcription, and technical requirements. See our Dataset Licensing & Compliance and Dataset Delivery & Security pages for details on rights documentation, packaging, and secure transfer.

Need the Full Dataset?

Request access to the full 4K and HD English speech video and audio dataset or discuss a custom audiovisual dataset requirement with Wavebreak Media.

Selected Partners

Selected Wavebreak Media partners