RESOURCE GLOSSARY

AI Training Data Glossary

Buyer-focused definitions for sourcing, evaluating, licensing, preparing, and delivering commercial AI training datasets, covering dataset types, annotations, metadata, provenance, releases, quality, multimodal alignment, and procurement.

How to Use This Glossary

Use these definitions when preparing dataset requirements, reviewing sample data, or comparing terminology across technical, legal, procurement, and delivery workflows.

Definitions describe how Wavebreak Media uses these terms in dataset sourcing and procurement. They do not replace project-specific legal, technical, or contractual review.

For purchasing guidance, see the Dataset Buyer's Guide and its AI Dataset Quality Checklist.

Last reviewed: August 2026. Maintained by Wavebreak Media.

Core AI Training Data Terms

AI training data
Data used to train or fine-tune an AI model. Related validation, evaluation, and benchmark datasets are used to tune or measure model performance and should remain separate from the training data.
Dataset
An organized collection of files and supporting information used for a specific AI, machine learning, evaluation, or data analysis workflow.
Generative AI training data
Data used to train, fine-tune, or evaluate models that generate or transform text, images, video, audio, designs, or other content. See Generative AI Training Data.
Training set
The portion of a dataset used to teach a model patterns, features, categories, relationships, or behaviors.
Validation set
A held-out portion of data used during model development to select settings, compare model versions, and detect overfitting. It remains separate from both the training and final test sets.
Test set
A separate portion of data used to evaluate final model performance after training and tuning, without overlapping the training data.
Fine-tuning data
Data used to adapt an existing model to a specific task, domain, behavior, format, or customer requirement. It is often more targeted than the data used for initial model training.
Evaluation dataset
Data kept separate from training and used to test how well a model performs on defined tasks, categories, conditions, or edge cases. See Model Evaluation Datasets.
Benchmark dataset
A fixed or versioned evaluation dataset used with defined tasks, reference answers, and metrics to compare model performance across versions, vendors, or technical approaches.

Dataset Type Terms

Image dataset
A collection of still images used for tasks such as classification, object recognition, visual search, computer vision, multimodal AI, or model evaluation. See Image Datasets for AI Training.
Video dataset
A collection of video files, clips, or extracted frames used for video understanding, motion analysis, activity recognition, generative video, or temporal AI tasks. See Video Datasets for AI Training.
Audio dataset
A collection of speech, environmental sounds, object sounds, music, or other audio recordings prepared for audio classification, generation, transcription, source separation, or model evaluation.
Audio-visual dataset
A multimodal dataset that connects video with synchronized audio using shared files, timestamps, event boundaries, or alignment metadata.
Document dataset
A collection of PDFs, scanned documents, forms, reports, or other document files used for OCR, extraction, classification, retrieval, layout understanding, or Document AI. See Document Datasets for AI Training.
Template dataset
A collection of editable design layouts, presentation templates, or structured creative assets used for design automation, layout understanding, component prediction, or generative design workflows. See Template Datasets.
Text dataset
A collection of text-first data used for NLP, classification, extraction, retrieval, summarization, instruction tuning, or language-model tasks, distinct from document datasets in its focus on language content rather than page layout. See Text Datasets for AI Training.
Multimodal dataset
A dataset that links two or more modalities, such as synchronized video-audio, image-text, video-text, or document-text data, using shared identifiers, timestamps, captions, or other alignment fields. See Multimodal AI Datasets.
Image-text dataset
A dataset that connects images with captions, descriptions, labels, questions, answers, or other text fields for vision-language training, retrieval, captioning, or evaluation.
Video-text dataset
A dataset that connects video clips or frames with captions, descriptions, questions, answers, or metadata for video understanding, retrieval, captioning, or vision-language tasks.
Human-authored data
Text, documents, captions, instructions, answers, or other records created by people against defined subject, style, factual, structural, and quality requirements rather than generated automatically.
Synthetic data
Data generated artificially to represent required patterns, events, conditions, or edge cases rather than captured directly from the target real-world event. Its generation method and intended use should be documented.
Custom dataset
A dataset produced or curated around a buyer's specific model task, content coverage, format, metadata, annotation, quality, or licensing requirements. See Custom Dataset Creation.

Computer Vision and Visual AI Terms

Computer vision
An AI field focused on helping models interpret images, videos, scenes, objects, people, actions, and visual patterns. See Computer Vision Datasets.
Object recognition
An AI task focused on identifying or classifying objects, products, tools, devices, or other visual entities in images or video. See Object Recognition Datasets.
Object detection
The task of identifying and locating objects within an image or video frame, commonly using class labels and bounding boxes or other spatial annotations.
Image classification
The task of assigning one or more predefined category labels to an image.
Image segmentation
The task of dividing an image or frame into labeled regions or object instances, commonly using pixel-level or polygon-based masks.
Scene understanding
The task of interpreting the broader context of an image or video, including its environment, objects, people, activities, conditions, and visual relationships.
Human activity recognition
The task of identifying actions, movement, gestures, routines, or behavior categories in images, videos, or frame sequences. See Human Activity Recognition Datasets.
Object tracking
The task of maintaining an object's identity and location across consecutive video frames, commonly using frame-level boxes, masks, points, or trajectories.
Facial expression analysis
The task of analyzing visible facial expression categories or face-centered visual cues without treating them as definitive evidence of a person's internal emotional state. See Facial Expression Analysis Datasets.
Avatar and presenter AI workflows
AI workflows involving visual data for virtual presenters, talking-head systems, or presenter-style footage, including gesture, posture, framing, gaze, and on-camera delivery. See Avatar & Presenter AI Datasets.

Language, Multimodal, and VLM Terms

Vision-language model (VLM)
An AI model that connects visual inputs with language information such as captions, descriptions, questions, answers, or instructions. See Vision-Language Model Datasets.
VLM dataset
A vision-language dataset prepared for tasks such as captioning, visual question answering, cross-modal retrieval, image-text understanding, video-text understanding, or model evaluation.
Captioned dataset
A dataset of images, videos, audio, documents, or other media paired with descriptive text, usually linked through stable identifiers or file-level records.
Visual question answering
A task where a model receives visual content and a related question, then produces or selects an answer based on the visible information.
Cross-modal retrieval
The task of finding relevant content across different modalities, such as retrieving images with text queries or matching video clips to descriptions.
Natural language processing (NLP)
AI workflows focused on language tasks such as classification, extraction, summarization, retrieval, and text understanding, distinct from Document AI, which also involves visual layout and OCR.
Large language model dataset
Text or multimodal data prepared for LLM pretraining, fine-tuning, instruction following, retrieval, domain adaptation, safety testing, or model evaluation.
Instruction-tuning data
Examples that pair an instruction or prompt with a target response so a language or multimodal model can learn to follow task-specific directions. The examples may also include context, constraints, rationale fields, or quality labels.
Retrieval corpus for RAG
A collection of documents, passages, records, or other source material made searchable for retrieval-augmented generation. A retrieval corpus supports grounding and retrieval but is not automatically model-training data.

Document AI and OCR Terms

Optical character recognition (OCR)
The process of detecting and extracting text from document images, scanned pages, photographs, screenshots, or other visual files.
Document AI
AI workflows that read, classify, extract, retrieve, or understand information from documents. See Document Datasets for AI Training.
Document classification
The task of assigning a document to one or more predefined classes based on its text, layout, visual appearance, metadata, or combined features.
Form extraction
The task of identifying and extracting structured fields, values, selections, or relationships from forms.
Table extraction
The task of detecting and extracting rows, columns, cells, headers, values, and relationships from tables in documents.
Layout understanding
The task of interpreting document structure, including text blocks, tables, headers, footers, forms, sections, fields, and visual hierarchy.
Reading order
The intended sequence in which text blocks, fields, tables, captions, and other document regions should be read or processed.
Extracted text
Text obtained from a document, image, or video frame through OCR, manual transcription, document parsing, or another extraction workflow.

Annotation and Metadata Terms

Data annotation
Task-specific reference information added to data for model training or evaluation, such as class labels, bounding boxes, masks, keypoints, temporal boundaries, captions, or extracted fields.
Metadata
Descriptive, technical, administrative, or rights information about dataset files, such as stable IDs, formats, capture attributes, provenance, release status, and version data. Metadata is not necessarily a model-training label.
Dataset label
A controlled class, category, state, or target assigned to a file, object, frame, action, document, field, or other dataset element for a defined model task.
Caption
A text description connected to an image, video, frame, audio file, document, or other dataset item.
Keyword tag
A short descriptive term used to organize, filter, or search dataset content. A keyword tag is not necessarily a controlled class label or ground-truth annotation.
Taxonomy
A structured category system used to organize classes, labels, tags, or metadata fields, often through parent-child relationships.
Ontology
A formal representation of entities, properties, relationships, and rules used to describe how concepts in a dataset relate to one another.
Ground truth
A verified or accepted reference answer, label, annotation, or measurement used to train a model or assess its outputs.
Bounding box
A rectangular annotation marking the location and approximate extent of an object or region within an image or video frame.
Segmentation mask
A pixel-level or region-based annotation identifying the area occupied by a class, object instance, foreground subject, or other target region.
Keypoint annotation
A set of labeled points marking defined locations such as body joints, facial landmarks, object parts, or other reference positions.
Temporal annotation
An annotation that identifies when an event, action, state, sound, shot, or other segment begins and ends within time-based data.
Quality assurance (QA)
The process of reviewing dataset files, labels, metadata, documentation, and delivery structure against defined acceptance criteria for consistency and usability.

Licensing, Rights, and Trust Terms

Rights-cleared dataset
A dataset for which the provider has documented the rights and permissions supporting the specified licensed use. Rights-cleared does not mean unrestricted use; permitted AI activities, releases, limitations, and license terms must still be reviewed.
Dataset licensing
The contractual framework defining how a dataset may be used, including permitted AI activities, commercial scope, restrictions, documentation, and delivery terms. See Dataset Licensing & Compliance.
Commercial AI use
Using data in workflows connected to paid products, enterprise systems, commercial model development, customer-facing services, or other business applications, subject to the agreed license.
Model release
Documentation recording a person's permission for specified uses of their identifiable likeness. Whether a release covers a particular AI training or commercial use depends on its language and the accompanying dataset license. See Model & Property Releases.
Property release
Documentation granting specified permissions relating to identifiable private property, interiors, locations, artwork, or other protected assets. Its applicability to a particular AI use must be reviewed with the dataset license.
Likeness rights
Rights and legal interests relating to the use of a person's identifiable appearance, image, voice, or representation. Applicable requirements vary by jurisdiction and intended use.
Usage scope
The activities a buyer is permitted to perform with a dataset under an agreed license, such as training, fine-tuning, evaluation, benchmarking, internal research, or commercial product development.
Dataset provenance
Information showing where data came from, how it was sourced, created, or transformed, and what records support source traceability. See Dataset Provenance & Source Traceability.
Chain of rights
The documented sequence showing how relevant rights or licensing authority pass from creators or rights holders through the provider to the permitted buyer use.
Data exclusivity
A contractual restriction on licensing the same data to other parties. Its scope may be limited by dataset, use case, customer, industry, geography, model type, or time period.

Dataset Delivery and Operations Terms

Dataset delivery
The process of packaging, transferring, documenting, and providing dataset files to the buyer. See Dataset Delivery & Secure Transfer.
Secure dataset delivery
An agreed workflow for transferring and handling dataset files using appropriate access controls, transfer methods, integrity checks, and buyer-specific security requirements.
Dataset package
The organized delivery set containing the agreed data files and, where applicable, metadata, annotations, manifests, documentation, release information, provenance records, and usage terms.
Dataset manifest
A structured record of files included in a dataset, commonly containing stable identifiers, paths, formats, sizes, split assignments, versions, checksums, or related metadata references.
Checksum
A value calculated from a file and used to verify that the file was transferred or stored without unintended alteration. A checksum verifies integrity but does not itself provide access control or encryption.
Dataset versioning
The practice of assigning identifiable versions and recording changes as dataset files, annotations, metadata, splits, manifests, or documentation are updated.
Sample dataset
A smaller set of representative files and supporting information provided for buyer review before full licensing, purchase, or production, subject to the applicable sample terms.

Dataset Quality Terms

Dataset quality
The degree to which a dataset meets its intended model task and acceptance criteria across relevance, coverage, technical integrity, consistency, and annotation accuracy. Licensing, provenance, releases, and delivery affect commercial suitability but require separate review. See the AI Dataset Quality Checklist.
Dataset coverage
How well a dataset represents the required categories, scenarios, subjects, objects, actions, environments, formats, conditions, or edge cases.
Dataset diversity and variation
Useful variation across relevant data dimensions, such as subjects, objects, scenes, environments, layouts, viewpoints, lighting, formats, languages, or capture conditions.
Dataset representativeness
The degree to which dataset content reflects the populations, environments, conditions, categories, and cases expected in the intended deployment or evaluation context.
Class balance
The distribution of examples across labels or categories. Appropriate balance depends on the model task and target environment and does not always require equal class sizes.
Dataset deduplication
The process of identifying and managing exact duplicates, near-duplicates, derivative files, repeated scenes, or other forms of unintended repetition within or across dataset splits.
Data leakage and train-test contamination
Unintended overlap or information transfer between training and evaluation data, including shared files, near-duplicates, subjects, scenes, sessions, or derived assets that can make performance results unreliable.
Annotation accuracy
The degree to which labels or annotations match the agreed ground truth, ontology, instructions, and acceptance criteria.
Hard negative
A non-target example that closely resembles the target class or commonly causes model errors, included to improve discrimination or test model robustness.
Dataset noise
Irrelevant, corrupted, ambiguous, low-quality, or incorrectly labeled data that reduces usefulness for the intended model task.
Edge case
An unusual, difficult, or less common example that remains relevant to the intended model behavior or deployment environment.

Buyer and Procurement Terms

Dataset brief
A written description of a buyer's dataset requirements, including model task, data type, content coverage, volume, formats, metadata, annotations, quality, licensing, and delivery. See Custom Dataset Creation.
Dataset requirements
The technical, content, metadata, annotation, quality, rights, licensing, security, and delivery conditions a dataset must meet.
Enterprise AI dataset
A dataset prepared for organizational or commercial use with the documentation, procurement, legal, security, licensing, quality, and delivery requirements needed for enterprise review. See Enterprise AI Datasets.
Dataset vendor
A company or supplier that licenses, curates, produces, annotates, packages, or delivers datasets for AI or machine-learning workflows.
Procurement review
The buyer-side process of assessing a dataset and vendor across requirements such as pricing, licensing, provenance, documentation, security, delivery, commercial suitability, and contractual terms.

Use These Terms to Evaluate or Request Data

Move from definitions to dataset evaluation, available collections, or a project-specific request.

Evaluate a Dataset

Review the questions, specifications, licensing checks, and quality criteria buyers should assess before approving training data.

Dataset Buyer's Guide

Browse Available Datasets

Review representative licensed image, video, audio, motion, and multimodal collections available from Wavebreak Media.

Browse Dataset Library

Request a Custom Dataset

Submit the required modality, content coverage, volume, metadata, annotations, rights, quality criteria, and delivery requirements.

Request Dataset

Selected Partners

Selected Wavebreak Media partners