VIDEO DATASET

Large-Scale HD AI-Generated Video Dataset for AI Training

Access 162,810 AI-generated horizontal video clips delivered through a single Full HD 1920 × 1080 profile. With approximately 242.5 hours of synthetic content, the collection provides a large standardized corpus for synthetic-video detection, representation learning, retrieval, classification and model evaluation.

The dataset is designed for workloads where training volume and consistent input geometry are more important than resolution diversity. Unlike the separate 4K + HD AI-generated collection, this corpus removes mixed-resolution variation and provides a substantially larger pool of standardized synthetic examples for training, validation and held-out evaluation. Asset-level records can also document synthetic status, generation workflow, provenance, generation date, prompt availability, reference-input status, human review and known generation limitations where retained.

162,810
Clips
242.5
Total Hours
5.4s
Average Clip Length
Full HD
Format
Horizontal
Orientation
AI-Generated
Content Type

Dataset Preview

Representative previews from the standardized horizontal HD AI-generated corpus, including varied synthetic subjects, environments, actions, visual styles, and generation artifacts.

The examples illustrate content and artifact variation within a consistent HD format, making it easier to compare generated scenes without resolution differences.

Key Highlights

  • • 162,810 AI-generated clips for large training, validation and evaluation splits
  • • Single 1920 × 1080 horizontal delivery profile across the corpus
  • • Explicit synthetic-content status for positive-class construction
  • • Consistent frame geometry for controlled training and model comparison
  • • Generation workflow, provenance, prompt and reference-input status available where retained
  • • Human-review state and known generation limitations can support QA and error analysis

Metadata Fields

clip_id Unique identifier linking the Full HD file to its manifest record
title Descriptive title of the AI-generated video clip
description Natural-language description of the visible generated content
keywords Keyword tags describing subjects, environments, objects, visual concepts, and scenes in the clip
resolution Video resolution; Full HD 1920 × 1080 in this dataset
format_conformance Validation status for the standardized horizontal Full HD delivery profile
duration Length of the individual video clip
fps Frames per second
orientation Video orientation; horizontal in this dataset
generation_workflow Generation route and subsequent Full HD packaging or normalization status where retained
provenance Available production, ownership, ingestion, and transformation history for the synthetic clip
synthetic_content Explicit boolean or controlled-value flag identifying the clip as AI-generated
generation_date Generation or production date where retained
model_or_production_source Generating model, platform, studio, or production source where known and permitted for disclosure
prompt_availability Availability level for the prompt: full, partial, summary, unavailable, or restricted
reference_input_status Whether generation used an image, video, or other reference input and whether that input can be supplied
human_review Recorded human review for technical conformity, visible artifacts, labeling, or policy checks
known_limitations Known generation issues such as identity drift, temporal inconsistency, object persistence, text legibility, anatomy, or physics errors

Technical Specifications

Dataset Type

Standardized Synthetic Video: HD Horizontal

Primary Technical Intent

Scalable classification, retrieval, synthetic-video detection, and efficient model training

Clip Count

162,810 clips

Total Duration

242.5 hours

Total Duration in Minutes

Approximately 14,548.1 minutes

Average Clip Length

Approximately 5.4 seconds

Resolution

1920 × 1080

Format

Full HD

Codec

H.264

Bitrate

Approximately 22 Mbps

Container

MP4

Orientation

Horizontal

Aspect Ratio

16:9

Format Consistency

Single 1920 × 1080 delivery profile for predictable decode and batching

Synthetic-Content Label

Explicit asset-level flag suitable for classification targets

Generation Provenance

Workflow, source, and date fields supplied where retained and permitted

Prompt & Reference Status

Availability values recorded without implying that source inputs are always deliverable

Human QA

Review status and known limitation notes supplied where recorded

Licensing, Documentation, and Delivery

Wavebreak Media can provide applicable licensing, generation provenance, and disclosure status information for the standardized HD AI-generated video content, with dataset delivery organized around agreed file, metadata, and technical requirements. See our Dataset Licensing & Compliance and Dataset Delivery & Security pages for details on rights documentation, packaging, and secure transfer.

Need the Full Dataset?

Request a representative Full HD synthetic-video sample or a larger subset based on training volume, metadata requirements, generation provenance, prompt status, review state and known limitation categories.

Selected Partners

Selected Wavebreak Media partners