RESOURCE GLOSSARY

AI Training Data Glossary

A practical glossary of AI training data and dataset terminology for teams buying, evaluating, licensing, and preparing datasets for commercial AI workflows.

Use this glossary to understand common terms related to dataset licensing, provenance, annotation, metadata, model releases, multimodal datasets, OCR, computer vision, Vision-Language Models, custom dataset creation, and secure dataset delivery.

How to Use This Glossary

Use these definitions when preparing dataset requirements, reviewing sample data, or comparing terminology across technical, legal, procurement, and delivery workflows.

For broader purchasing guidance, see the Dataset Buyer's Guide and its AI Dataset Quality Checklist.

Core AI Training Data Terms

AI training data
The data used to train, fine-tune, evaluate, benchmark, or enrich an AI model, including images, videos, documents, text, templates, metadata, labels, captions, or multimodal data.
Dataset
An organized collection of files and supporting information used for a specific AI, machine learning, evaluation, or data analysis workflow.
Training set
The portion of a dataset used to teach a model patterns, features, categories, relationships, or behaviors.
Validation set
A portion of data used during model development to tune settings, compare model versions, and reduce overfitting, kept separate from the training set.
Test set
A separate portion of data used to evaluate final model performance after training and tuning, without overlapping the training data.
Fine-tuning data
Data used to adapt an existing model to a specific task, domain, format, or customer requirement, typically smaller than the initial training dataset.
Evaluation dataset
Data used to test how well a model performs on defined tasks, categories, or edge cases, kept separate from training data.
Benchmark dataset
A dataset used to compare model performance across versions, vendors, systems, or model approaches.

Dataset Type Terms

Image dataset
A collection of still images used for tasks such as classification, object recognition, visual search, computer vision, multimodal AI, or model evaluation. See Image Datasets for AI Training.
Video dataset
A collection of video files, clips, or extracted frames used for video understanding, motion analysis, activity recognition, or temporal AI tasks. See Video Datasets for AI Training.
Document dataset
A collection of PDFs, scanned documents, forms, reports, or other document files used for OCR, extraction, classification, retrieval, or document AI. See Document Datasets for AI Training.
Template dataset
A collection of design layouts, presentation templates, or editable creative assets used for design automation, layout understanding, or creative AI workflows. See Template Datasets.
Text dataset
A collection of text-first data used for NLP, classification, extraction, search, retrieval, summarization, or language model tasks, distinct from document datasets in its focus on language content rather than layout.
Multimodal dataset
A dataset that connects more than one data type, such as image-text pairs, video-text pairs, captions, or structured fields, used when a model needs to understand relationships across visual and language data. See Multimodal AI Datasets.
Image-text dataset
A dataset that connects images with captions, descriptions, labels, or other text fields, supporting Vision-Language Models, captioning, and retrieval tasks.
Video-text dataset
A dataset that connects video clips or frames with captions, descriptions, or metadata fields, supporting video understanding, VLM workflows, and captioning tasks.
Custom dataset
A dataset built or curated around a buyer’s specific model task, content category, format, or licensing scope, used when available datasets do not match the required use case. See Custom Dataset Creation.

Computer Vision and Visual AI Terms

Computer vision
An AI field focused on helping models interpret images, videos, scenes, objects, people, actions, and visual patterns.
Object recognition
An AI task focused on identifying or classifying objects, products, tools, devices, or other visual entities in images or video. See Object Recognition Datasets.
Object detection
The task of locating objects within an image or video frame, often using bounding boxes, regions, or other spatial annotations, usually requiring more detailed annotation than simple classification.
Image classification
The task of assigning one or more predefined category labels to an image.
Scene understanding
The task of interpreting the broader context of an image or video, including environment, objects, people, activity, and visual relationships.
Human activity recognition
The task of identifying actions, movement, gestures, routines, or behavior categories in images, videos, or frame sequences. See Human Activity Recognition Datasets.
Facial expression analysis
The task of analyzing visible facial expression categories or face-centered visual cues, without inferring internal emotional state. See Facial Expression Analysis Datasets.
Avatar and presenter AI workflows
AI workflows involving visual data for virtual presenters, talking-head systems, or presenter-style footage, covering gesture, posture, framing, and on-camera delivery. See Avatar & Presenter AI Datasets.

Language, Multimodal, and VLM Terms

Vision-Language Model
An AI model that connects visual inputs with language-based information such as captions, descriptions, questions, answers, or instructions. See Vision-Language Model Datasets.
VLM dataset
A dataset designed for Vision-Language Model tasks such as captioning, visual question answering, retrieval, image-text understanding, or video-text understanding.
Captioned dataset
A dataset of images, videos, or other media paired with descriptive text.
Visual question answering
A task where a model answers questions about visual content.
Cross-modal retrieval
The task of finding relevant content across different data types, such as retrieving images using text queries or matching video clips to text descriptions.
Natural language processing, or NLP,
AI workflows focused on language tasks such as classification, extraction, summarization, retrieval, and text understanding, distinct from document AI, which also involves layout and OCR.

Document AI and OCR Terms

OCR, or optical character recognition,
The process of detecting and extracting text from images, scanned documents, or document files.
Document AI
AI workflows that read, classify, extract, retrieve, or understand information from documents. See Document Datasets for AI Training.
Form extraction
The task of identifying and extracting structured fields from forms.
Table extraction
The task of detecting and extracting rows, columns, cells, or other structured information from tables in documents.
Layout understanding
The task of interpreting the structure of a document, including text blocks, tables, headers, footers, forms, sections, and visual hierarchy.
Extracted text
Text pulled from a document, image, or video frame using OCR, manual extraction, or another processing workflow.

Annotation and Metadata Terms

Annotation
The process of adding structured information to dataset files, such as labels, captions, bounding boxes, categories, OCR fields, tags, or descriptions.
Metadata
Information that describes dataset files, such as category, caption, description, file type, label, tag, release note, provenance field, or usage note.
Label
A category or tag assigned to a file, object, frame, document, action, expression, or other dataset element.
Caption
A text description connected to an image, video, frame, or other media file.
Keyword tag
A short descriptive term used to classify or search dataset content.
Taxonomy
The structured category system used to organize labels, tags, classes, or metadata fields.
Ground truth
The reference answer or label used to train or evaluate a model.
QA, or quality assurance,
The process of reviewing dataset files, labels, metadata, documentation, and delivery structure for consistency and usability.

Licensing, Rights, and Trust Terms

Rights-cleared dataset
A dataset prepared with licensing, usage scope, source rights, release requirements, and documentation review in mind. Rights-cleared does not mean unlimited use — buyers should still review the license terms.
Dataset licensing
Defines how a dataset can be used, including permitted use, commercial scope, restrictions, and delivery terms. See Dataset Licensing & Compliance.
Commercial AI use
Using data in workflows connected to paid products, enterprise systems, commercial model development, or other business applications.
Model release
Documentation of permissions for using a recognizable person’s likeness in visual media, especially relevant for human-centered image and video datasets. See Model & Property Releases.
Property release
Documentation of permissions for using identifiable private property, interiors, locations, branded spaces, or other protected assets in visual media. See Model & Property Releases.
Likeness rights
Rights relating to the use of a person’s identifiable appearance, image, or representation.
Usage scope
What a buyer is allowed to do with a dataset under an agreed license, such as training, evaluation, fine-tuning, or commercial product development.
Dataset provenance
Information about where data came from, how it was sourced or produced, and what documentation supports source traceability. See Dataset Provenance & Source Traceability.

Dataset Delivery and Operations Terms

Dataset delivery
The process of packaging, transferring, documenting, and providing dataset files to the buyer. See Dataset Delivery & Secure Transfer.
Secure dataset delivery
An agreed workflow for transferring and handling dataset files using appropriate access controls, delivery methods, and buyer-specific transfer requirements.
Dataset package
The final organized delivery set that may include files, metadata, labels, documentation, release notes, provenance notes, and usage information.
File inventory
A list of files included in a dataset delivery package.
Versioning
The practice of tracking dataset versions as files, metadata, labels, or documentation change.
Sample dataset
A smaller subset of files provided for buyer review before purchase, licensing, or full production.

Dataset Quality Terms

Dataset quality
How well a dataset fits the intended model task, including relevance, consistency, file quality, metadata, licensing, and delivery structure. See the AI Dataset Quality Checklist.
Coverage
How well a dataset represents the required categories, scenarios, subjects, objects, formats, or contexts.
Diversity
Useful variation within a dataset, such as different scenes, environments, subjects, objects, layouts, camera angles, or document types.
Duplicate control
The process of identifying and managing repeated or near-repeated files within a dataset.
Noise
Irrelevant, low-quality, mislabeled, corrupted, or distracting data that reduces dataset usefulness.
Edge case
An unusual or less common example that may still be important for model performance.

Buyer and Procurement Terms

Dataset brief
A written description of a buyer’s dataset requirements, including use case, data type, categories, volume, metadata, licensing, and delivery format.
Dataset requirements
The technical, content, licensing, metadata, quality, and delivery conditions a dataset must meet.
Enterprise AI dataset
A dataset prepared for business or organizational use, often with procurement, legal, security, licensing, and documentation requirements. See Enterprise AI Datasets.
Dataset vendor
A company or supplier that provides, licenses, curates, produces, or delivers datasets for AI workflows.
Procurement review
The buyer-side process of checking vendor terms, pricing, licensing, documentation, delivery, and commercial suitability.

Related Resources

After reviewing glossary terms, use these resources to evaluate or request a dataset: Dataset Buyer's Guide, AI Training Datasets, Dataset Licensing & Compliance, Dataset Delivery & Secure Transfer, and Custom Dataset Creation.

Selected Partners

Selected Wavebreak Media partners