VIDEO + AUDIO DATASET

Everyday Object Sounds Dataset

Access 25,000 training-ready video clips with synchronized high-quality audio capturing everyday objects and materials producing isolated sound events. The collection provides approximately 156 hours of purpose-built audiovisual data for AI model development, sound synthesis, foley generation, audio event detection, and video-to-audio generation.

Each clip establishes a clear relationship between a visible action and its resulting recorded sound, reducing ambiguity in multimodal training. The dataset covers materials including wood, metal, glass, plastic, rubber, leather, rope, and other everyday objects, with structured metadata designed for precise audio-visual alignment and retrieval. Because every sound is tied to a visible action, the collection fits naturally into broader multimodal datasets built around cross-modal grounding.

25,000
Clips
~156
Total Hours
~22.5s
Average Clip Length
4K UHD
Video Format
48kHz Stereo
Audio
Horizontal
Orientation

Dataset Preview

Representative previews showing everyday objects, materials, and physical interactions paired with synchronized recorded sound events.

The examples illustrate how visible actions correspond to object and material sounds, supporting review of event timing, source identity, and audio-visual alignment.

Key Highlights

  • Isolated object and material sound events with clear visual causes
  • Structured metadata covering sound category, object, action, and resulting sound
  • Purpose-built for audio-visual AI and video-to-audio model development

Metadata Fields

filename Source filename or unique asset identifier
title Descriptive title of the object sound event
sound_category Material or sound category such as wood, metal, glass, plastic, rubber, leather, or rope
object Object producing or participating in the sound event
action Visible action responsible for producing the sound
resulting_sound Description or classification of the resulting recorded sound

Technical Specifications

Dataset Type

Synchronized Video + Object Sound Dataset

Content Type

Everyday objects and materials producing isolated sound events

Clip Count

25,000 clips

Total Duration

Approximately 156 hours

Average Clip Length

Approximately 22.5 seconds

Video Resolution

4K UHD

Orientation

Horizontal

Video Container

MOV / MP4

Frame Rate

24–30 fps

Audio Format

WAV

Audio Sample Rate

48kHz

Audio Channels

Stereo

Synchronization

Audio synchronized with corresponding visible action

Audio Content

Isolated sound effects from visible objects and material interactions

Spoken Language

None

Licensing, Documentation, and Delivery

Wavebreak Media can provide applicable licensing, provenance, and model or property release information for the synchronized video and object-sound content, with dataset delivery organized around agreed AI training and project requirements. See our Dataset Licensing & Compliance and Dataset Delivery & Security pages for details on rights documentation, packaging, and secure transfer.

Need the Full Dataset?

Request access to the Everyday Object Sounds Dataset or discuss custom object categories, sound events, metadata, licensing, and delivery requirements with Wavebreak Media.

Selected Partners

Selected Wavebreak Media partners