4K & HD Object & Foley Sounds Video + Audio Dataset for AI Training
Access a large-scale audiovisual dataset containing 25,657 horizontal video clips paired with recorded object and foley sounds, delivered in both 4K/UHD and HD. The collection provides approximately 93.9 hours of synchronized video and non-speech audio for AI training, fine-tuning, evaluation, enrichment, and commercial model-development workflows.
The dataset includes recordings of objects, materials, tools, and physical interactions, with examples such as drums, angle grinding, and sounds involving metal, glass, wood, plastic, and ceramic objects. Structured metadata supports audiovisual understanding, retrieval, classification, filtering, and other multimodal AI workflows. The corpus can be licensed alongside our other audio-visual datasets where sound and image must be learned together.
Dataset Preview
Representative previews showing visible object and material interactions paired with synchronized recorded foley and environmental sound.
The examples illustrate the relationship between physical actions and their corresponding sound events, including tools, surfaces, objects, and material interactions.
Key Highlights
- • Video clips are paired with recorded object and foley sounds for synchronized audiovisual learning
- • Non-speech recordings of objects, tools, materials, and physical interactions
- • Examples include drums, angle grinding, metal, glass, wood, plastic, and ceramic objects
- • Structured metadata for filtering, retrieval, classification, and AI workflow support
Metadata Fields
Example Metadata Record
title: Metal object interaction with recorded sound
description: Horizontal video showing a metal object interaction accompanied by synchronized non-speech audio
keywords: metal, object sound, foley, interaction, recorded audio, audiovisual, horizontal video
resolution: 3840x2160
duration: 00:00:13
fps: Available in source metadata
orientation: Horizontal
shoot_date: Available where applicable
model_release: Available where applicable
property_release: Available where applicable
Technical Specifications
Dataset Type
Video + Audio Dataset: Object & Foley Sounds Horizontal
Content Type
Video clips paired with recorded non-speech object and foley sounds
Clip Count
25,657 clips
Total Duration
93.9 hours
Total Duration in Minutes
Approximately 5,634.5 minutes
Average Clip Length
Approximately 13.2 seconds
4K/UHD Resolution
3840–4096 × 2160
HD Resolution
1920 × 1080
Formats
4K/UHD + HD
Codec
H.264
4K/UHD Bitrate
Approximately 87–93 Mbps
HD Bitrate
Approximately 22 Mbps
Container
MP4
Orientation
Horizontal
Aspect Ratio
16:9
Spoken Language
N/A, non-speech audio
Audio Content
Recorded object, material, tool, and foley sounds
Subtitles / Scripts
N/A, non-speech audio; no subtitles or scripts
Licensing, Documentation, and Delivery
Wavebreak Media can provide applicable licensing, provenance, and model or property release information for the audiovisual content, with dataset delivery organized around agreed video, audio, metadata, and technical requirements. See our Dataset Licensing & Compliance and Dataset Delivery & Security pages for details on rights documentation, packaging, and secure transfer.
Need the Full Dataset?
Request access to the full 4K and HD object and foley sounds video and audio dataset or discuss a custom audiovisual dataset requirement with Wavebreak Media.
Related Datasets
Everyday Object Sounds Dataset
Access 25,000 training-ready 4K video clips with synchronized object and material sounds for audio-visual AI, foley generation, sound synthesis, and multimodal learning.
View Dataset VIDEO + AUDIO DATASET4K & HD English Speech Video + Audio Dataset for AI Training
Access 1,126 English speech video clips with recorded audio in 4K and HD for multimodal AI training, speech recognition, audiovisual understanding, and evaluation.
View Dataset MATERIAL HANDLING EQUIPMENT VISUALSForklift Image Dataset
Access approximately 1,000 forklift images for computer vision, object recognition, industrial-equipment understanding, visual classification, retrieval, and multimodal AI workflows.
View Dataset
