VIDEO DATASET

HD AI-Generated Video Dataset for AI Training

Access 162,810 horizontal AI-generated clips delivered in a consistent Full HD 1920 × 1080 format. With approximately 242.5 hours of synthetic video, the corpus is designed for high-volume classification, retrieval, synthetic-video detection, embedding development, and efficient training pipelines that benefit from standardized decoding.

The metadata can pair uniform technical delivery with asset-level generation records: workflow, provenance, synthetic-content flag, generation date, model or production source where disclosure is permitted, prompt availability, reference-input status, human review, and known limitations. Its uniform delivery makes the corpus straightforward to combine with other generative AI training data in a single training pipeline.

162,810
Clips
242.5
Total Hours
5.4s
Average Clip Length
Full HD
Format
Horizontal
Orientation
AI-Generated
Content Type

Dataset Preview

Representative previews from the standardized horizontal HD AI-generated corpus, including varied synthetic subjects, environments, actions, visual styles, and generation artifacts.

The examples illustrate content and artifact variation within a consistent HD format, making it easier to compare generated scenes without resolution differences.

Key Highlights

  • Consistent Full HD delivery supports predictable decoding and batching
  • Large corpus suited to class-balanced sampling, retrieval indexing, and representation learning
  • Explicit synthetic-content flag for positive-class construction
  • Generation workflow, date, provenance, and permitted production-source status can be included
  • Prompt and reference-input availability recorded independently from the media file
  • Human-review state and known generation limitations can support QA and error analysis

Metadata Fields

clip_id Unique identifier linking the Full HD file to its manifest record
title Descriptive title of the AI-generated video clip
description Natural-language description of the visible generated content
keywords Keyword tags describing subjects, environments, objects, visual concepts, and scenes in the clip
resolution Video resolution; Full HD 1920 × 1080 in this dataset
format_conformance Validation status for the standardized horizontal Full HD delivery profile
duration Length of the individual video clip
fps Frames per second
orientation Video orientation; horizontal in this dataset
generation_workflow Generation route and subsequent Full HD packaging or normalization status where retained
provenance Available production, ownership, ingestion, and transformation history for the synthetic clip
synthetic_content Explicit boolean or controlled-value flag identifying the clip as AI-generated
generation_date Generation or production date where retained
model_or_production_source Generating model, platform, studio, or production source where known and permitted for disclosure
prompt_availability Availability level for the prompt: full, partial, summary, unavailable, or restricted
reference_input_status Whether generation used an image, video, or other reference input and whether that input can be supplied
human_review Recorded human review for technical conformity, visible artifacts, labeling, or policy checks
known_limitations Known generation issues such as identity drift, temporal inconsistency, object persistence, text legibility, anatomy, or physics errors

Technical Specifications

Dataset Type

Standardized Synthetic Video: HD Horizontal

Primary Technical Intent

Scalable classification, retrieval, synthetic-video detection, and efficient model training

Clip Count

162,810 clips

Total Duration

242.5 hours

Total Duration in Minutes

Approximately 14,548.1 minutes

Average Clip Length

Approximately 5.4 seconds

Resolution

1920 × 1080

Format

Full HD

Codec

H.264

Bitrate

Approximately 22 Mbps

Container

MP4

Orientation

Horizontal

Aspect Ratio

16:9

Format Consistency

Single 1920 × 1080 delivery profile for predictable decode and batching

Synthetic-Content Label

Explicit asset-level flag suitable for classification targets

Generation Provenance

Workflow, source, and date fields supplied where retained and permitted

Prompt & Reference Status

Availability values recorded without implying that source inputs are always deliverable

Human QA

Review status and known limitation notes supplied where recorded

Licensing, Documentation, and Delivery

Wavebreak Media can provide applicable licensing, generation provenance, and disclosure status information for the standardized HD AI-generated video content, with dataset delivery organized around agreed file, metadata, and technical requirements. See our Dataset Licensing & Compliance and Dataset Delivery & Security pages for details on rights documentation, packaging, and secure transfer.

Need the Full Dataset?

Request the standardized HD synthetic video corpus with the required sample size, class schema, provenance fields, prompt and reference status, review state, and delivery manifest.

Selected Partners

Selected Wavebreak Media partners