Talking Head, Avatar & Presenter AI Datasets
License camera-facing presenter video from Wavebreak Media's owned archive, or commission controlled recordings for talking-head and avatar AI. We produce the source footage and can scope scripts, speech recording and annotations around your model requirements.
Featured Datasets
Explore sample licensed datasets across image, video, audio, and text collections.








Avatar and Presenter Data by Task
The focus is complete presenter performances and speech-related motion. See facial expression datasets for isolated expression classes.
Talking Head Generation
Face-centered speaking sequences for speech-driven facial animation and portrait animation, with footage for temporal-consistency testing and output evaluation.
Lip-Sync and Visual Speech
Presenter video with synchronized source audio, where suitable and licensed, for audio-driven avatars, lip-sync, visual speech and dubbing evaluation.
Head Pose and Gaze
Head rotation, nods, gaze direction, blinking and mouth movement across speaking and non-speaking sequences, with the relevant facial dynamics visible.
Co-Speech Gestures
Speech with hand and body gestures, emphasis, holds, retractions and posture shifts. See human activity datasets for broader movement tasks.
Presenter Framing
Face, upper-body or full-body views for testing backgrounds, camera distances, aspect ratios and safe areas, with foreground segmentation scoped where needed.
Evaluation and Coverage
Held-out participants, sessions, scripts, poses and capture conditions for model evaluation, using agreed references and comparison criteria.
Presenter Performance and Capture Coverage
Archive availability is checked against the required delivery style and sequence depth. Specific scripts, languages, phonetic coverage, clean audio and balanced participant quotas may require custom recording.
Face-Centered Video
Close-ups and head-and-shoulders footage showing mouth, eyes and head movement. Face-Focused Portrait Images adds still inputs and portrait evaluation material.
Seated and Standing Delivery
Direct-to-camera monologues, interviews and demonstrations in medium, upper-body or full-body compositions, with seated or standing positions defined in the brief.
Scripted and Natural Delivery
Prepared lines, explanations, demonstrations, conversation and natural speech. Define speaking pace, emphasis and delivery style, including repeated takes where required.
Speech and Listening States
Speaking and non-speaking intervals, neutral holds, pauses, breathing and listening behavior. Specify the transitions and sequence depth needed across complete takes.
Capture Conditions
Camera angle, distance, resolution, frame rate and aspect ratio, plus lighting, background, wardrobe and audio setup. Select conditions that match the intended deployment.
Participant Variation
Face shape, hair, facial hair, glasses, clothing and presentation style. Define visible participant attributes and coverage targets that matter to your model requirements.
Audio, Video and Annotation Specification
Dataset Relationships
Stable participant, session, take, clip, frame, script and source-audio IDs. Preserve one-to-many relationships so each derivative can be traced to its recording.
Speech and Timing
Source audio, transcripts, utterance boundaries, words, phonemes, visemes and mouth states. Availability and alignment timestamps are confirmed or commissioned per project.
Visual Annotations
Face and body boxes, landmarks, head pose, gaze, facial motion, hand or body keypoints and foreground masks. Confirm annotation methods, framing fields and quality criteria.
Packaging and Delivery
Manifests, codecs, technical metadata and checksums where required. Agree output formats, quality status and secure dataset delivery before preparing the final package.
Archive Footage and Custom Presenter Production
Select a presenter subset from our video datasets or review multimodal datasets when paired speech and video are required. Assess samples before committing to annotation.
Wavebreak Media has produced professional visual media since 2005. Custom presenter production can cover requirements that available footage cannot meet, with capture, releases and acceptance criteria agreed before recording.
Frequently Asked Questions (FAQ)
Not always. Check synchronization, speech clarity, noise, music, file format and applicable rights in a sample. Clean, controlled speech recording may require new production.
Specify transcript format, utterance boundaries, word or phoneme timing, viseme taxonomy and acceptable alignment error. Agree review criteria on representative samples before scaling preparation.
The standard offering is 2D image and video data. Scans, meshes, depth, rigging, blendshapes and motion-capture outputs require a separate scope and feasibility review; they are not included by default.
Yes, with suitable source video and audio, transcripts, translated scripts or new recordings. Confirm language and speaker availability, timing references and evaluation criteria for the intended dubbing task.
Where participant and session identities are known or controlled during capture. Define grouping for related takes, scripts, adjacent frames and derived clips. Unknown archive relationships can limit separation checks.
No blanket permission is implied. Likeness-, identity-, biometric-, or voice-specific uses require explicit project scope and rights review. Dataset licensing and compliance requirements, permitted uses, and restrictions are defined for the project.
Request an Avatar or Presenter AI Dataset
Send your model task, presenter format, speech and language needs, sequence length and approximate volume. Include a script or sample recording brief if available.

