LICENSED IMAGE-TEXT AND VIDEO-TEXT DATA

Vision-Language Model Datasets for AI Training

License image-text and video-text data from Wavebreak Media's wholly owned visual archive for captioning, visual question answering, retrieval, grounding, instruction tuning, and evaluation, or commission custom visual capture and language annotation.

Datasets can be curated by visual category, pair type, caption depth, question-answer format, temporal alignment, grounding schema, language, hard negatives, split rules, usage rights, and delivery format.

Vision-Language Data by Task

Configure vision-language model datasets around the required visual-language relationship, reference text, grounding depth, and evaluation structure.

Image-Text Pretraining and Retrieval

Image-text pairs linked by captions, descriptions, tags, or concepts, with positive and negative relationships, relevance labels, and query-gallery splits where specified.

Image Captioning and Description

Short captions, dense descriptions, or multiple references per image written to agreed content, detail, style, and groundedness rules.

Video Captioning and Temporal Understanding

Video-text pairs aligned at clip or segment level with actions, events, scene changes, and temporal order using timestamps or frame references where required.

Visual Question Answering (VQA)

Image- or video-grounded questions and answers covering objects, actions, counts, spatial relationships, visible text, scenes, or temporal events.

Visual Grounding and Referring Expressions

Examples that teach a model to associate referring phrases with the correct object, region, frame, or moment in the visual source.

Visual Instruction Tuning and Evaluation

Single- or multi-turn visual prompts, responses, held-out examples, hard negatives, and controlled criteria for supervised fine-tuning or evaluation.

Model Evaluation Datasets
Vision-language dataset workflow linking images and video with captions, questions, answers, grounding references, and retrieval pairs

Visual Content Available for VLM Data

Wavebreak Media can assess archive coverage and custom production across the following visual domains. Availability is confirmed against the required subject depth, pair volume, text format, and annotation scope.

People and Social Interaction

Portraits, conversation, communication, family activity, teamwork, and other individual, interpersonal, or group scenes. The People and Human Interaction Images Dataset provides a licensed still-image source for captioning, VQA, retrieval, and social-scene descriptions.

Human Activities and Object Use

Daily routines, exercise, mobility, device use, tool handling, and people interacting with products or equipment.

Products, Retail, and Food

Consumer goods, packaged products, household objects, ingredients, meals, shopping activity, and items shown alone or in context.

Workplaces, Business, and Technology

Professional environments, meetings, computer use, devices, communication, collaboration, and task-based business scenes.

Travel, Cities, and Outdoor Scenes

Transport, streets, landmarks, mobility, tourism, nature, weather, outdoor activity, and environmental context.

Animation, Motion, and Creative Assets

Animation, motion graphics, designed compositions, and other non-photographic visual content where suitable source assets are available.

Pairing, Annotation, and Dataset Splits

Selected visual and language data can be prepared with:

  • Pair structure - stable visual, clip, frame, text, question, and answer IDs; one-to-one or one-to-many relationships; pair type; and source relationships.
  • Language fields - captions, dense descriptions, questions, answers, answer choices, rationales, instructions, responses, language or locale, and authoring or review status where specified.
  • Grounding and alignment - object or region references, boxes, masks, frame IDs, timestamps, temporal segments, and referring expressions where commissioned.
  • Quality, splits, and delivery - groundedness and consistency checks, hard-negative validation, near-duplicate and leakage controls, source-aware grouping, manifests, agreed schemas, checksums where required, and secure dataset delivery.

Vision-Language Dataset Options

Existing visual metadata, archive curation, custom visual production, and project-specific language annotation can be used separately or combined under one specification.

Captioned Image and Image-Text Data

Select still-image collections for captioning, VQA, retrieval, or region grounding, then assess whether existing text is suitable or new annotation is required.

Image Datasets for AI Training

Captioned Video and Video-Text Data

Select video collections for clip or segment descriptions, temporal VQA, or frame grounding while preserving source sequence and timing relationships.

Video Datasets for AI Training

Custom VLM Dataset Production

Commission visual capture, language annotation, pairing, grounding, quality review, and delivery for a project-defined VLM schema and acceptance criteria.

Custom Dataset Creation

Why Wavebreak Media for Vision-Language Data?

Since 2005, Wavebreak Media has produced and managed an extensive image archive and more than one million wholly owned video assets. This provides visual coverage across people, activities, products, objects, environments, and professionally produced sequences.

Projects can combine eligible visual assets with project-specific language annotation, pairing schemas, quality checks, and licensing for the intended AI use. Available provenance and applicable releases can be reviewed through the dataset licensing and compliance process.

Frequently Asked Questions (FAQ)

Not automatically. Existing assets may include descriptions, titles, keywords, categories, or technical metadata. These fields are assessed, normalized, expanded, or replaced according to the project specification.

Yes, when scoped. Projects can define author qualifications, caption depth, grounding rules, question and answer formats, prohibited inferences, review stages, and acceptance thresholds.

Yes, subject to the required languages, locales, authoring method, and reviewer availability. Translation and original native-language annotation should be specified separately.

VLM datasets specifically connect images or video with language. Multimodal datasets can combine a broader set of modalities, including synchronized video and audio, sensor data, or other paired inputs.

Yes. Splits can keep related source assets, shoots, sequences, extracted frames, near-duplicate visuals, and derived text together so they do not leak across training and evaluation sets.

Licensing is finalized for selected visual and language data and the intended AI use. Available provenance and applicable releases are reviewed before approval; permitted uses are defined in the agreement.

Request a Vision-Language Model Dataset

Send the target visual domains, VLM task, image or video requirements, text fields, language, grounding or temporal alignment, volume, split rules, rights requirements, and delivery schema. Wavebreak Media will assess archive-curated and custom production options against the specification.

Selected Partners

Selected Wavebreak Media partners