Long-Horizon Human Procedural Video Dataset for AI Training
Access a growing collection of purpose-built 4K videos capturing complete real-world human tasks from beginning to end. The dataset currently contains 579 completed videos and approximately 19 hours of content, with sequences averaging around two minutes rather than the short isolated actions common in conventional activity-recognition datasets.
Each sequence preserves the progression of a task across time: the starting situation, successive human actions, repeated interaction with tools or objects, intermediate changes and the visible outcome of the activity. This continuous context allows models to observe how individual actions depend on earlier steps and contribute to completion of a longer procedure.
Current coverage spans 15 task categories including food preparation, cooking, baking, cleaning, DIY, tool use, crafts, organisation and packing, gardening, office tasks, health and wellness, personal care, workshop activities and everyday object interaction.
Dataset Preview
Representative frames from longer 4K sequences covering complete human tasks across household, workplace, workshop, wellness and outdoor environments. Each source video preserves actions and object interactions as part of a continuous procedure rather than as isolated clips.
The source sequences preserve temporal continuity from the beginning of an activity through intermediate actions and object-state changes to a visible task outcome, providing context that individual short action clips cannot capture.
Key Highlights
- • Complete human tasks captured from beginning to visible outcome
- • Multi-step action sequences preserved within continuous videos
- • Human-object interaction maintained across successive task stages
- • Approximately two-minute average sequences for longer temporal context
- • 15 task categories covering household, workshop, office, outdoor and personal activities
- • Professionally produced 4K footage with rights-cleared participants
Metadata Fields
Technical Specifications
Dataset Type
Long-Form Human Procedural Video Dataset
Current Library Size
579 completed videos
Production Status
Additional production ongoing
Current Total Duration
Approximately 19 hours
Average Clip Length
Approximately 2 minutes
Video Resolution
4K
Production Quality
Professional commercial production quality
Production Standard
Consistent across the dataset
Content Structure
Complete multi-step activities preserving task progression from beginning to outcome
Interaction Focus
Sequential human actions, repeated object manipulation, intermediate task states, and complete workflow context
Audio
48kHz WAV stereo audio
Participant Rights
Rights-cleared participants
Content Categories
Food Preparation; Cooking; Baking; Household Tasks; DIY & Home Improvement; Tool Usage; Craft Activities; Cleaning; Organisation & Packing; Gardening; Office Tasks; Health & Wellness; Personal Care; Workshop Activities; Everyday Object Interaction
Licensing, Documentation, and Delivery
Wavebreak Media can provide applicable licensing, provenance, and model or property release information for the long-form procedural video, with dataset delivery organized around agreed project and delivery requirements. See our Dataset Licensing & Compliance and Dataset Delivery & Security pages for details on rights documentation, packaging, and secure transfer.
Need the Full Dataset?
Request representative complete sequences from specific task categories to evaluate temporal continuity, human-object interaction, workflow complexity, metadata and production quality, or discuss additional custom procedural scenarios for future capture.
Related Datasets
Egocentric First-Person POV Household Activities Dataset
Access purpose-built first-person household activity video for embodied AI, robotics, object interaction, procedural understanding, and egocentric vision model development.
View Dataset CONTROLLED HUMAN MOTION CAPTUREHuman Motion Dataset (DV01)
Access about 6,000 purpose-built human motion videos with structured metadata for action recognition, gesture understanding, behaviour modelling, video understanding, and multimodal AI.
View Dataset 3-WAY SYNCHRONIZED CAPTURESynchronized Multi-View Video Dataset
Access synchronized 4K multi-camera video captured from three viewpoints for multi-view synthesis, camera-aware generation, shot planning, cinematic understanding, and video editing AI.
View Dataset
