Talking Head, Avatar & Presenter AI Datasets
License camera-facing presenter video from Wavebreak Media's wholly owned archive for talking head and avatar systems, lip-sync evaluation, facial motion, and co-speech gestures, or commission controlled recordings for a defined model specification.
Datasets can be curated by framing, duration, participant, session, head pose, gaze, gesture, background, source audio, annotations, split rules, usage rights, and delivery format.
Featured Datasets
Explore sample licensed datasets across image, video, audio, and text collections.
Avatar and Presenter Data by Task
Configure archive-curated or custom-produced video around sustained camera-facing delivery, speech-related motion, presenter framing, and evaluation requirements. Unlike facial expression datasets, this offering focuses on complete presenter performances rather than isolated expression classes.
Talking Head Generation and Evaluation
Face-centered speaking sequences for speech-driven facial animation, portrait animation, temporal-consistency testing, and talking-head output evaluation.
Audio-Visual Speech and Lip-Sync
Presenter video with synchronized source audio where technically suitable and licensed, supporting audio-driven avatar, lip-sync, and visual speech workflows.
Head Pose, Gaze, and Facial Motion
Video sequences covering head rotation, nods, gaze direction, blinking, mouth movement, and other visible facial dynamics across time.
Facial Expression DatasetsCo-Speech Gesture and Body Motion
Temporal relationships between speech and hand or body gestures, including onset, emphasis, holds, retractions, posture shifts, and pauses where labeled.
Human Activity Recognition DatasetsPresenter Framing and Compositing
Consistent face, upper-body, or full-body framing across backgrounds, camera distances, aspect ratios, and foreground-segmentation requirements.
Evaluation and Coverage Expansion
Held-out participants, sessions, scripts, poses, gestures, backgrounds, and difficult capture conditions selected against defined evaluation criteria.
Model Evaluation DatasetsPresenter Performance and Capture Coverage
Wavebreak Media can assess archive and custom production options against the required presenter style, framing, movement, setting, and sequence depth.
Specific scripts, languages, speaking rates, repeated takes, controlled phonetic coverage, clean audio, and balanced participant coverage generally require custom production.
Face-Centered Talking Head Video
Close-up and head-and-shoulders sequences with the face, mouth, eyes, and natural head movement visible throughout the take. For still-image inputs and portrait-based evaluation, review the Face-Focused Portrait Images Dataset.
Seated and Standing Presenters
Direct-to-camera monologues, demonstrations, interviews, and explanatory delivery in medium, upper-body, or full-body compositions.
Scripted and Natural Delivery
Prepared lines, direct-to-camera presentation, explanation, demonstration, conversation, and natural speaking behavior where available or produced.
Speech, Pauses, and Listening States
Speaking and non-speaking intervals, neutral holds, pauses, breathing, listening behavior, and transitions between delivery states.
Camera and Environment Variation
Framing, camera angle, distance, resolution, frame rate, aspect ratio, background, lighting, wardrobe, and recording environment.
Participant and Appearance Variation
Project-defined participant coverage, face shape, hair, facial hair, glasses, clothing, presentation style, and visible appearance attributes.
Audio, Video, Annotations, and Dataset Splits
Selected presenter assets can be prepared with:
- Dataset relationships - stable participant, session, take, clip, frame, script, and source-audio IDs, with one-to-many relationships preserved where specified.
- Speech and timing data - synchronized audio, transcripts, utterance boundaries, words, phonemes, visemes, mouth states, and alignment timestamps where available or commissioned.
- Visual annotations - face and body boxes, landmarks, head pose, gaze, facial motion, hand or body keypoints, foreground masks, framing, and quality status where commissioned.
- Split and delivery controls - participant-, session-, script-, and shoot-aware grouping, adjacent-frame and near-duplicate rules, manifests, codecs, technical metadata, checksums where required, and secure dataset delivery.
Avatar and Presenter Dataset Options
Archive curation, synchronized audio-visual subsets, custom presenter recording, and project-specific annotation can be used separately or combined under one specification.
Archive-Curated Presenter Video
Build a project-specific collection from eligible footage selected for camera-facing delivery, framing, movement, setting, technical quality, and available metadata.
Video Datasets for AI TrainingSynchronized Audio-Visual Speech Data
Assess eligible speaking footage for usable synchronized audio, then define transcript, alignment, annotation, and quality requirements for the selected subset.
Multimodal AI DatasetsCustom Scripted Presenter Production
Commission authorized participants, scripts, languages, delivery styles, framing, audio capture, repetitions, annotations, releases, and acceptance criteria under a controlled brief.
Custom Dataset CreationWhy Wavebreak Media for Presenter and Avatar Data?
Since 2005, Wavebreak Media has produced and managed an extensive image archive and more than one million wholly owned video assets. Its human-centered footage includes camera-facing talent, business presentations, product demonstrations, lifestyle scenes, exercise, and social interaction across professional capture conditions.
Projects can combine eligible archive footage with controlled recording, project-defined annotations, quality checks, and licensing for the intended AI use. Available provenance and applicable model or property releases can be reviewed through the dataset licensing and compliance process.
Frequently Asked Questions (FAQ)
Not always. Source footage must be checked for synchronized audio, speech quality, background noise, music, format, and applicable rights. Custom recording is preferable when controlled audio is required.
Yes, when commissioned. Projects can specify transcript format, utterance boundaries, word or phoneme timing, viseme taxonomy, alignment tolerances, and review criteria.
The standard offering is 2D image and video data. 3D scans, meshes, depth, rigging, blendshapes, or motion-capture outputs are not assumed and require an explicitly scoped custom project.
Yes, when the project includes suitable source video and audio, transcripts, translated scripts, or new recordings. Languages, speakers, timing references, and evaluation criteria must be specified.
Yes, when participant and session relationships are available or controlled during production. Related takes, scripts, adjacent frames, and derived clips can remain within one split to reduce leakage.
No blanket permission is implied. Likeness-, identity-, biometric-, or voice-specific uses require explicit project scope and rights review. Permitted uses and restrictions are defined in the agreement.
Request an Avatar or Presenter AI Dataset
Send the model task, presenter format, participant criteria, scripts or speech requirements, language, framing, gestures, audio needs, sequence length, volume, annotations, split rules, rights requirements, and delivery format. Wavebreak Media will assess archive-curated, audio-visual, and custom production options against the specification.

