CLEAN SFX + EVENT ONSET ANNOTATIONS

Video-to-Audio Foley Dataset with Event Onset Annotations

Access approximately 25,500 video-to-audio clips with visible object-level sound events paired with synchronized clean SFX/Foley audio. The dataset provides approximately 70 finished hours across two existing collections, generally with one isolated visible sound event per clip and a corresponding WAV file.

Audio is limited to the relevant sound effect, with no speech, music or unrelated ambient audio. Event labels, clip descriptions, per-clip FCPXML files and decimal-second onset timestamps can be supplied, with a target onset accuracy within 20 ms for prepared data.

~25,500
Video + SFX Clips
~70 hrs
Finished Content
≤20 ms
Target Onset Accuracy
48 kHz
WAV Audio
5,500+
24-bit WAV Available
No Speech
Clean SFX / Foley

Dataset Preview

Representative samples show visible object-level sound events paired with synchronized clean Foley/SFX audio. Prepared evaluation clips include event labels, onset timing, source metadata and per-clip FCPXML for review of audiovisual correspondence and temporal alignment.

Video-to-audio Foley samples with visible object-level sound events and synchronized clean SFX for audiovisual alignment and model evaluation.

Key Highlights

  • Approximately 25,500 visible object-level sound-event clips
  • Clean synchronized SFX/Foley audio paired with corresponding video
  • Generally one isolated visible sound event per clip
  • No speech, music or unrelated background audio
  • Event-onset timestamps with target accuracy within 20 ms
  • Individual FCPXML export available for each prepared clip
  • At least 5,500 clips available with clean lossless 24-bit / 48 kHz WAV where available from original projects
  • Event labels, clip descriptions and source metadata available for structured AI workflows

Metadata Fields

clip_id Unique identifier for the prepared audiovisual clip
event_label Label describing the visible sound event
clip_description Description derived from existing captions or source descriptions and checked during preparation
onset_seconds Sound-event onset represented as a decimal-seconds value
onset_marker Named marker identifying the onset or transient in the source project
video_filename Filename of the delivered video asset
audio_filename Filename of the corresponding synchronized WAV asset
fcpxml_filename Individual FCPXML export associated with the prepared clip
source_collection Source collection associated with the clip
source_metadata Available technical and source information retained from the original asset or project

Technical Specifications

Dataset Type

Video-to-Audio Foley / SFX Dataset

Content Type

Visible object-level sound events with synchronized clean SFX/Foley audio

Total Volume

Approximately 25,500 clips

Total Duration

Approximately 70 finished hours

Higher-Quality Collection

Approximately 5,500 clips / 15 finished hours

Additional Collection

Approximately 20,000 clips / 55 finished hours

Sound Events per Clip

Generally one isolated visible event

Preferred Video Delivery

MOV

Additional Video Source Format

Original edited projects include BRAW source video where applicable

Additional Collection Format

Primarily compressed MP4; technical specifications verified during preparation

Audio Format

WAV

Audio Sample Rate

48 kHz

Audio Channels

Stereo

24-bit Audio Availability

At least 5,500 clips can be supplied with clean lossless 24-bit / 48 kHz WAV where available from original projects

Existing-Library Audio Baseline

Synchronized uncompressed 48 kHz stereo WAV; bit depth confirmed during preparation where 24-bit delivery is required

Audio Content

Relevant SFX/Foley only

Speech

None

Music

None

Unrelated Ambient Audio

Excluded from clean prepared audio

Language

Language-independent

Subtitles

Not applicable

Scripts / Transcripts

Not applicable

Event Labels

Available with prepared data

Clip Descriptions

Available with prepared data

Onset Annotation

Decimal-seconds timestamp + named FCPXML marker

Target Onset Accuracy

Within 20 ms

FCPXML

One individual export per prepared clip

Metadata Delivery

CSV + FCPXML + available source metadata

Licensing, Documentation, and Delivery

Wavebreak Media can provide applicable licensing, provenance, and model or property release information for the footage, with dataset delivery organized around agreed file, metadata, and technical requirements. See our Dataset Licensing & Compliance and Dataset Delivery & Security pages for details on rights documentation, packaging, and secure transfer.

Need Video-to-Audio Training Data with Precise Foley Timing?

Request access to the Video-to-Audio Foley Dataset or provide your required sound-event categories, clip volume, audio quality, onset accuracy, metadata and licensing specifications for a prepared or custom collection.

Selected Partners

Selected Wavebreak Media partners