Video-to-Audio Foley Dataset with Event Onset Annotations
Access approximately 25,500 video-to-audio clips with visible object-level sound events paired with synchronized clean SFX/Foley audio. The dataset provides approximately 70 finished hours across two existing collections, generally with one isolated visible sound event per clip and a corresponding WAV file.
Audio is limited to the relevant sound effect, with no speech, music or unrelated ambient audio. Event labels, clip descriptions, per-clip FCPXML files and decimal-second onset timestamps can be supplied, with a target onset accuracy within 20 ms for prepared data.
Dataset Preview
Representative samples show visible object-level sound events paired with synchronized clean Foley/SFX audio. Prepared evaluation clips include event labels, onset timing, source metadata and per-clip FCPXML for review of audiovisual correspondence and temporal alignment.
Video-to-audio Foley samples with visible object-level sound events and synchronized clean SFX for audiovisual alignment and model evaluation.
Key Highlights
- • Approximately 25,500 visible object-level sound-event clips
- • Clean synchronized SFX/Foley audio paired with corresponding video
- • Generally one isolated visible sound event per clip
- • No speech, music or unrelated background audio
- • Event-onset timestamps with target accuracy within 20 ms
- • Individual FCPXML export available for each prepared clip
- • At least 5,500 clips available with clean lossless 24-bit / 48 kHz WAV where available from original projects
- • Event labels, clip descriptions and source metadata available for structured AI workflows
Metadata Fields
Technical Specifications
Dataset Type
Video-to-Audio Foley / SFX Dataset
Content Type
Visible object-level sound events with synchronized clean SFX/Foley audio
Total Volume
Approximately 25,500 clips
Total Duration
Approximately 70 finished hours
Higher-Quality Collection
Approximately 5,500 clips / 15 finished hours
Additional Collection
Approximately 20,000 clips / 55 finished hours
Sound Events per Clip
Generally one isolated visible event
Preferred Video Delivery
MOV
Additional Video Source Format
Original edited projects include BRAW source video where applicable
Additional Collection Format
Primarily compressed MP4; technical specifications verified during preparation
Audio Format
WAV
Audio Sample Rate
48 kHz
Audio Channels
Stereo
24-bit Audio Availability
At least 5,500 clips can be supplied with clean lossless 24-bit / 48 kHz WAV where available from original projects
Existing-Library Audio Baseline
Synchronized uncompressed 48 kHz stereo WAV; bit depth confirmed during preparation where 24-bit delivery is required
Audio Content
Relevant SFX/Foley only
Speech
None
Music
None
Unrelated Ambient Audio
Excluded from clean prepared audio
Language
Language-independent
Subtitles
Not applicable
Scripts / Transcripts
Not applicable
Event Labels
Available with prepared data
Clip Descriptions
Available with prepared data
Onset Annotation
Decimal-seconds timestamp + named FCPXML marker
Target Onset Accuracy
Within 20 ms
FCPXML
One individual export per prepared clip
Metadata Delivery
CSV + FCPXML + available source metadata
Licensing, Documentation, and Delivery
Wavebreak Media can provide applicable licensing, provenance, and model or property release information for the footage, with dataset delivery organized around agreed file, metadata, and technical requirements. See our Dataset Licensing & Compliance and Dataset Delivery & Security pages for details on rights documentation, packaging, and secure transfer.
Need Video-to-Audio Training Data with Precise Foley Timing?
Request access to the Video-to-Audio Foley Dataset or provide your required sound-event categories, clip volume, audio quality, onset accuracy, metadata and licensing specifications for a prepared or custom collection.
Related Datasets
4K & HD Object & Foley Sounds Video + Audio Dataset for AI Training
Access 25,657 video clips with recorded object and foley sounds in 4K and HD for multimodal AI training, audiovisual understanding, sound recognition, and evaluation.
View Dataset VIDEO + AUDIO DATASETEveryday Object Sounds Dataset
Access 25,000 training-ready 4K video clips with synchronized object and material sounds for audio-visual AI, foley generation, sound synthesis, and multimodal learning.
View Dataset VIDEO + AUDIO DATASET4K & HD English Speech Video + Audio Dataset for AI Training
Access 1,126 English speech video clips with recorded audio in 4K and HD for multimodal AI training, speech recognition, audiovisual understanding, and evaluation.
View Dataset
