HUMAN-ANNOTATED VIDEO DATASET

Camera Motion Annotation Dataset

Access 100,000 professionally curated video clips with human-generated, frame-accurate camera motion annotations. The collection provides approximately 417 hours of annotated 4K video designed for training and evaluating models that recognize, classify, and understand camera movement patterns across diverse real-world scenes.

Each clip includes quality-controlled temporal annotations identifying camera motion type with precise start and end boundaries. Structured metadata includes motion labels, frame rate, resolution, content category, and natural-language scene descriptions, providing supervised data for video understanding, temporal segmentation, scene analysis, retrieval, and motion-aware generation. Motion-labelled video of this kind is used alongside our broader computer vision datasets.

100,000
Annotated Clips
~417
Total Hours
~15s
Average Clip Length
4K
Resolution
24 fps
Frame Rate
Frame-Accurate
Annotations

Dataset Preview

Representative video examples showing source scenes and camera movements together with corresponding frame-accurate temporal annotations.

The examples illustrate how human-generated motion labels align with the source video, including the temporal boundaries of individual camera movements.

Key Highlights

  • Frame-accurate camera motion start and end boundaries
  • Human-generated, quality-controlled camera motion labels
  • Clean source video without graphics, captions, or overlays
  • Structured CSV annotations aligned to individual video assets
  • Rights-cleared professionally produced video content

Metadata Fields

clip_id Identifier linking the annotation data to the corresponding video clip
camera_motion_label Human-generated label identifying the annotated camera movement type
start_time Frame-accurate temporal start boundary for the annotated motion
end_time Frame-accurate temporal end boundary for the annotated motion
frame_rate Frame rate of the associated video asset
resolution Resolution of the associated video asset
content_category Content category associated with the video scene
scene_description Natural-language description of the visible scene

Technical Specifications

Dataset Type

Human-Annotated Camera Motion Video Dataset

Clip Count

100,000 annotated clips

Total Duration

Approximately 417 hours

Average Clip Length

Approximately 15 seconds

Video Resolution

4K

Frame Rate

24 fps

Annotation Type

Human-generated camera motion labels

Annotation Precision

Frame-accurate temporal start and end boundaries

Annotation Format

CSV

Annotation Alignment

Individual CSV annotation data aligned to corresponding video assets

Scene Metadata

Content category and natural-language scene descriptions

Source Quality

Professionally produced video

Source Presentation

Clean video without graphics, captions, or overlays

Audio

None

Licensing, Documentation, and Delivery

Wavebreak Media can provide applicable licensing, provenance, and model or property release information for the rights-cleared camera motion video and annotations, with dataset delivery organized around agreed model-development, benchmarking, and delivery requirements. See our Dataset Licensing & Compliance and Dataset Delivery & Security pages for details on rights documentation, packaging, and secure transfer.

Need the Full Dataset?

Request access to the Camera Motion Annotation Dataset or discuss annotation structure, motion labels, metadata, licensing, benchmarking, and technical delivery requirements with Wavebreak Media.

Selected Partners

Selected Wavebreak Media partners