VIDEO + AUDIO DATASET

4K & HD English Speech Video + Audio Dataset for AI Training

Access a video and audio dataset containing 1,126 horizontal clips with recorded English speech, delivered in both 4K/UHD and HD. The collection provides approximately 15.4 hours of synchronized audiovisual content for AI training, fine-tuning, evaluation, enrichment, and commercial model-development workflows.

The dataset combines visible human-centered video with spoken English audio, making it suitable for multimodal systems that need to interpret relationships between speech, speakers, facial and body cues, and visual context. Structured metadata supports retrieval, filtering, organization, and integration into AI development pipelines.

1,126
Clips
15.4
Total Hours
49.4s
Average Clip Length
4K + HD
Formats
Horizontal
Orientation
English
Spoken Language

Dataset Preview

Representative previews showing English speech video clips with synchronized recorded audio and varied human-centered visual contexts.

Preview assets shown here represent content available within the dataset.

Key Highlights

  • 1,126 horizontal video clips paired with recorded English speech audio
  • Approximately 15.4 total hours of synchronized audiovisual content
  • Average clip duration of approximately 49.4 seconds
  • Available in both 4K/UHD and HD formats
  • English spoken-language content
  • Human-centered video suitable for audiovisual model development
  • Structured metadata for filtering, retrieval, organization, and AI workflow support
  • No subtitles or scripts supplied; English audio can be transcribed on request

Metadata Fields

title Descriptive title of the video clip
description Natural-language description of the visible video content
keywords Keyword tags describing people, activities, environments, subjects, and concepts in the clip
resolution Video resolution, including 4K/UHD or HD
duration Length of the individual audiovisual clip
fps Frames per second
orientation Video orientation; horizontal in this dataset
shoot_date Date associated with the original recording where available
model_release Flag indicating available model-release status where applicable
property_release Flag indicating available property-release status where applicable

Example Metadata Record

title: English-speaking person on camera

description: Horizontal video showing a person speaking in English with synchronized recorded audio

keywords: person, speaking, English speech, video audio, human, conversation, horizontal video

resolution: 3840x2160

duration: 00:00:49

fps: Available in source metadata

orientation: Horizontal

shoot_date: Available where applicable

model_release: Available where applicable

property_release: Available where applicable

Technical Specifications

Dataset Type

Video + Audio Dataset — English Speech Horizontal

Content Type

Video clips paired with recorded English speech audio

Clip Count

1,126 clips

Total Duration

15.4 hours

Total Duration in Minutes

Approximately 926.2 minutes

Average Clip Length

Approximately 49.4 seconds

4K/UHD Resolution

3840–4096 × 2160

HD Resolution

1920 × 1080

Formats

4K/UHD + HD

Codec

H.264

4K/UHD Bitrate

Approximately 87–93 Mbps

HD Bitrate

Approximately 22 Mbps

Container

MP4

Orientation

Horizontal

Aspect Ratio

16:9

Spoken Language

English

Audio Content

Recorded speech audio paired with video

Subtitles / Scripts

N/A — no subtitles or scripts supplied

Transcription

English audio transcribable on request

Licensing, Documentation, and Delivery

Wavebreak Media can provide applicable licensing, provenance, and model or property release information for the audiovisual content, with dataset delivery organized around agreed video, audio, metadata, transcription, and technical requirements. See our Dataset Licensing & Compliance page for more detail.

Need the Full Dataset?

Request access to the full 4K and HD English speech video and audio dataset or discuss a custom audiovisual dataset requirement with Wavebreak Media.

Selected Partners

Selected Wavebreak Media partners