Multimodal AI Datasets for Video, Audio and Text
Wavebreak Media provides licensed multimodal AI datasets built around paired and synchronized video, audio, text, and structured metadata. Current collections include human speech video with recorded audio and object, material, tool, and foley video with synchronized sound.
License an existing dataset, request a project-specific subset of Wavebreak Media-owned content, or commission purpose-built data with defined alignment, transcription, metadata, provenance, release, licensing, and delivery requirements.
Featured Datasets
Explore sample licensed datasets across image, video, audio, and text collections.
What Is Multimodal AI Training Data?
Multimodal AI training data combines two or more distinct data modalities while preserving the relationships between them. Video can be paired with synchronized audio, while transcripts, captions, labels, and technical metadata can be linked to complete clips or specific timecodes.
For image-text and video-text data designed specifically for captioning, visual question answering, grounding, or retrieval, see Vision-Language Model Datasets.
Audio-Visual Datasets and Synchronized Modalities
An audio visual dataset connects video frames with recorded speech, object sounds, or other synchronized audio. Native audiovisual containers, shared asset IDs, timestamps or timecodes where available, manifests, and schema documentation keep each component reliably matched during training and evaluation.
Synchronized Speech Video
Pair video with native recorded speech audio while preserving source synchronization.
Object and Foley Sounds
Pair video with synchronized object, material, tool, or foley sounds.
Transcripts and Captions
Pair video and audio with transcripts or captions prepared on request.
Linked Audio Metadata
Connect audio with speaker, scene, content, and technical metadata.
Asset IDs and Manifests
Preserve media relationships through shared asset IDs, timecodes, and delivery manifests.
Custom Modality Schemas
Define custom modality combinations and relationship schemas for the target model workflow.

Multimodal AI Training and Evaluation Use Cases
Licensed multimodal datasets can support:
- Automatic and audiovisual speech recognition
- Speech and visual-context alignment
- Audiovisual event and sound-source recognition
- Speaker and speaking-activity detection
- Object interaction and sound recognition
- Multimodal model fine-tuning, evaluation, and benchmarking
Available Licensed Multimodal Datasets
English Speech Video + Audio
The 4K & HD English Speech Video + Audio Dataset includes 1,126 horizontal clips and approximately 15.4 hours of synchronized human speech video. Transcripts, timestamps, and metadata can be prepared to an agreed specification.
Object and Foley Video + Audio
The 4K & HD Object & Foley Sounds Video + Audio Dataset includes 25,657 synchronized non-speech clips covering objects, materials, tools, and physical interactions.
Everyday Object Sounds
The Everyday Object Sounds Dataset includes 25,000 4K clips and approximately 156 hours of recorded actions and corresponding object sounds.
Licensed, Curated and Custom Multimodal Data
Projects can license an existing collection, curate a project-specific subset using agreed selection criteria, or commission custom multimodal datasets from owned source media or new production. Commercial usage rights, provenance records, applicable model or property releases, and permitted model uses are reviewed through Wavebreak Media's dataset licensing and compliance process.
Frequently Asked Questions (FAQ)
Existing collections focus on synchronized video and audio, including speech, objects, materials, tools, and foley sounds. Curated or custom projects can also connect video and audio with transcripts, captions, speaker information, scene labels, technical metadata, or other agreed data fields.
Relationships can be preserved through native audiovisual containers, shared asset IDs, consistent filenames, timestamps or timecodes where available, manifests, and documented schemas. The exact alignment structure is defined according to the source material and model workflow.
Yes. Speech video can be prepared with clip-level or time-aligned transcripts, captions, timestamps, speaker fields, and agreed metadata. Language coverage, transcription conventions, accuracy requirements, and schema must be defined for the project.
Yes. Wavebreak Media can select a project-specific subset by subject, environment, technical format, clip duration, audio characteristics, release status, metadata availability, and other agreed criteria. Target volume and acceptance rules are confirmed before delivery.
Yes. Multimodal datasets can be licensed under dataset-specific commercial terms defining permitted training, fine-tuning, evaluation, deployment, and generated-output uses. Applicable provenance and model or property release documentation can be provided for review.
Yes. A custom project can define the required modalities, subjects, environments, volume, synchronization, transcription, metadata, file relationships, licensing scope, and delivery structure. Data can be sourced from Wavebreak Media-owned media or created through new production.
Request Multimodal AI Datasets
Provide the required modalities, model use case, content coverage, target volume, alignment or transcription needs, metadata, licensing scope, and delivery timeline. Wavebreak Media will identify relevant existing collections or define a custom multimodal dataset.

