Predict Grounding DINO

GPU Dockerfile
First run can take longer while the worker pulls the configured Docker image.

Overview

Runs Grounding DINO over a dataset and saves the detections as annotations in the output dataset.

Pick a Destination dataset to add the predictions to a dataset you already have, such as a labeling pool or an active-learning set: it keeps what it holds, and an image it already annotates gets the new prediction instead. Leave it empty and the node writes into a dataset of its own, one version per run.

The model detects objects from their names, so a public checkpoint finds your classes without any training. That makes it the fastest way to pre-label a fresh dataset.

The badge counts the annotated images and opens the output dataset.

Dataset Predict Grounding DINO Evaluate object detection
  • Grounding DINO Swin-T (lighter, ~170M params)
  • Grounding DINO Swin-B (~240M params)
  • Grounding DINO Swin-L (best quality, ~340M params)

Prerequisites

RequirementWhat you need
GPU workerA worker with a GPU.
Image datasetAn existing Dataset with images.
A model to runA model of this workspace or the zoo, picked in Model. Train Grounding DINO produces one, and the picker copies a zoo entry in.

How it works

1

Resolves the checkpoint from the selected source and loads it onto the GPU.

2

Creates the output dataset and adds the source images to it.

3

Builds the text prompt — the picked classes, or the noun phrases of the sentence — and pins the output dataset to it.

4

Runs detection over the images in batches, dropping detections below the confidence threshold.

5

Writes the predictions as annotations and shows how many images were annotated on the badge.

FAQ

No. Predictions land in the output dataset and the source keeps whatever it already had.

Replace leaves only the new predictions in the output dataset, which is what you want before an evaluation node. Keep both carries the input annotations over next to the predictions, which is what you want when comparing them side by side.

Yes, that is the point of the model. It matches objects against the class names it is given, so a public checkpoint already handles most everyday objects. Accuracy drops on domain-specific or visually unusual classes, which is when fine-tuning pays off.

They are what the model matches against, so wording matters. A plain, concrete noun works better than an internal code or an abbreviation.

A vocabulary is the usual choice: the classes are reusable, the output dataset is pinned to them, and a later run of the same flow detects the same things. A text prompt suits a one-off look at a dataset, or a description no class name captures — “a person wearing a yellow helmet”. Only the noun phrases of the sentence can be predicted, so anything it does not name is not detected. Typing short names separated by commas is treated as a plain class list, not as a sentence.

Run the matching evaluation node against a ground-truth dataset. Prediction alone reports how many images were annotated, not how correct the annotations are.

Yes, through Apply model. It reads the framework off the model asset and dispatches to this node, which is what you want when the model is only picked at runtime.

Inputs

Dataset
datasetRequired

Dataset to run detection on. Images are annotated when the dataset holds any, videos frame by frame otherwise. Key: DATASET.

Destination dataset
dataset

Dataset the predictions are added to, keeping what it already holds. An image it already annotates gets the new prediction instead. Leave empty to use the node’s own output dataset, which a rerun writes a new version of. Key: DST_DATASET.

Model
modelRequired

Model to run inference with, from this workspace or the zoo. The node is handed the resolved version and runs on its weights: the checkpoint stored for it, or the public weights of its architecture for a pretrained entry that hosts none. Key: MODEL.

What to detect
selectDefaults to vocabularyRequired

Grounding DINO is prompted with text, so what it looks for is chosen here rather than baked into the weights. A vocabulary asks for a list of classes. A text prompt grounds the noun phrases of a sentence. Key: PROMPT_MODE.

Options:

  • Class vocabulary (vocabulary)
  • Text prompt (prompt)
Classes
ontology

Vocabulary to detect - every class is asked for, and ticking a subset asks for that subset only. Tags are never asked for, since a tag has no box. Type a name to add a class the ontology does not have yet, or leave it empty to detect what a fine-tuned checkpoint was trained on. Visible when PROMPT_MODE is vocabulary. Key: CLASSES.

Text prompt
stringRequired

Sentence to ground, for example - The picture contains watermelon, flower, and a white bottle. Its noun phrases become the predicted classes, so a plain comma-separated list such as cat, dog, car is read as those three classes instead. Placeholder: There is a person wearing a yellow helmet.. Visible when PROMPT_MODE is prompt. Key: PROMPT.

Config
yamlRequired

Inference settings in YAML - how confident a detection must be to be kept, how many images go through the GPU at once, and the resolution they are run at. Set that resolution to the one the model was trained at, so training and inference see images the same way. Key: CONFIG.

Download batch size
integerDefaults to 64Required

Number of images fetched at a time while the GPU works through the ones already downloaded. Minimum: 1. Key: DOWNLOAD_BATCH_SIZE.

Prefetch batches
integerDefaults to 4Required

How far downloading may run ahead of inference, counted in batches. Raise it when a slow connection leaves the GPU idle, at the cost of disk space. Minimum: 1. Key: PREFETCH_BATCHES.

Outputs

Output Dataset
dataset

Dataset with the input images plus the Grounding DINO prediction annotations: the destination dataset when one is picked, otherwise the node’s own. Shown as an artifact. Key: OUTPUT_DATASET.

Predicted
string

Number of assets annotated. Shown on the node as a badge. Selecting the badge opens OUTPUT_DATASET. Key: PRED_COUNT.

Models and configuration

Model is the one place the weights come from: a model of this workspace, or one of the zoo entries listed above. The node is handed the resolved version and runs on its weights — the checkpoint stored for it, or the public weights of its architecture for an entry that hosts none.

What to detect decides where the prompt comes from, because the weights hold no class list of their own. With Class vocabulary the classes of the picked ontology are the prompt, and ticking a subset asks for that subset only. Leaving it empty falls back to the classes a fine-tuned checkpoint was trained on. With Text prompt the sentence is the prompt and its noun phrases become the predicted classes.

The output dataset is pinned to the vocabulary the predictions are named in, so the classes stay readable in the labeling tool and in evaluation.

Config holds the inference settings. conf_threshold is the score a detection needs to survive, and batch_size how many images the model sees at once.

Runtime

The node runs in its own GPU container. The first run takes longer while the worker pulls the image and downloads the selected weights.

Images are fetched ahead of the model to keep the GPU busy. Download Batch Size sets how many arrive at once and Prefetch Batches how many wait on local disk, so raise them for fast storage and lower them when disk is tight.

JSON config

Machine-readable node interface for automation and advanced usage.

{
"name": "Predict Grounding DINO",
"description": "Run Grounding DINO object detection inference on a dataset and write the images and the Grounding DINO prediction annotations into the output dataset.",
"category": "Predict",
"namespace": "ovalbee",
"templateKey": "models/grounding_dino/predict_grounding_dino",
"version": "v1",
"inputs": [
{
"key": "DATASET",
"label": "Dataset",
"type": "dataset",
"description": "Dataset to run detection on. Images are annotated when the dataset holds any, videos frame by frame otherwise.",
"required": true,
"default": null,
"visibleWhen": null,
"options": {
"creatable": false
}
},
{
"key": "DST_DATASET",
"label": "Destination dataset",
"type": "dataset",
"description": "Dataset the predictions are added to, keeping what it already holds. An image it already annotates gets the new prediction instead. Leave empty to use the node's own output dataset, which a rerun writes a new version of.",
"required": false,
"default": null,
"visibleWhen": null
},
{
"key": "MODEL",
"label": "Model",
"type": "model",
"description": "Model to run inference with, from this workspace or the zoo. The node is handed the resolved version and runs on its weights: the checkpoint stored for it, or the public weights of its architecture for a pretrained entry that hosts none.\n",
"required": true,
"default": null,
"visibleWhen": null,
"options": {
"framework": "grounding_dino",
"task_type": "object_detection"
}
},
{
"key": "PROMPT_MODE",
"label": "What to detect",
"type": "select",
"description": "Grounding DINO is prompted with text, so what it looks for is chosen here rather than baked into the weights. A vocabulary asks for a list of classes. A text prompt grounds the noun phrases of a sentence.\n",
"required": true,
"default": "vocabulary",
"visibleWhen": null,
"options": {
"options": [
{
"label": "Class vocabulary",
"value": "vocabulary"
},
{
"label": "Text prompt",
"value": "prompt"
}
]
}
},
{
"key": "CLASSES",
"label": "Classes",
"type": "ontology",
"description": "Vocabulary to detect - every class is asked for, and ticking a subset asks for that subset only. Tags are never asked for, since a tag has no box. Type a name to add a class the ontology does not have yet, or leave it empty to detect what a fine-tuned checkpoint was trained on.\n",
"required": false,
"default": null,
"visibleWhen": {
"key": "PROMPT_MODE",
"operator": "equals",
"value": "vocabulary"
},
"options": {
"selectable": true,
"creatable": true
}
},
{
"key": "PROMPT",
"label": "Text prompt",
"type": "string",
"description": "Sentence to ground, for example - The picture contains watermelon, flower, and a white bottle. Its noun phrases become the predicted classes, so a plain comma-separated list such as cat, dog, car is read as those three classes instead.\n",
"required": true,
"default": null,
"visibleWhen": {
"key": "PROMPT_MODE",
"operator": "equals",
"value": "prompt"
},
"options": {
"placeholder": "There is a person wearing a yellow helmet."
}
},
{
"key": "CONFIG",
"label": "Config",
"type": "yaml",
"description": "Inference settings in YAML - how confident a detection must be to be kept, how many images go through the GPU at once, and the resolution they are run at. Set that resolution to the one the model was trained at, so training and inference see images the same way.\n",
"required": true,
"default": "conf_threshold: 0.3\nbatch_size: 4\nimgsz: 800 # inference resolution (short-edge ref; long edge ~imgsz*1.667). Match train imgsz.\n",
"visibleWhen": null
},
{
"key": "DOWNLOAD_BATCH_SIZE",
"label": "Download batch size",
"type": "number",
"description": "Number of images fetched at a time while the GPU works through the ones already downloaded.",
"required": true,
"default": 64,
"visibleWhen": null,
"options": {
"type": "integer",
"min": 1
}
},
{
"key": "PREFETCH_BATCHES",
"label": "Prefetch batches",
"type": "number",
"description": "How far downloading may run ahead of inference, counted in batches. Raise it when a slow connection leaves the GPU idle, at the cost of disk space.",
"required": true,
"default": 4,
"visibleWhen": null,
"options": {
"type": "integer",
"min": 1
}
}
],
"outputs": [
{
"key": "OUTPUT_DATASET",
"label": "Output Dataset",
"type": "dataset",
"description": "Dataset with the input images plus the Grounding DINO prediction annotations: the destination dataset when one is picked, otherwise the node's own.",
"artifact": true,
"badge": false,
"badgeOpens": null,
"hidden": false
},
{
"key": "PRED_COUNT",
"label": "Predicted",
"type": "string",
"description": "Number of assets annotated.",
"artifact": false,
"badge": true,
"badgeOpens": "OUTPUT_DATASET",
"hidden": false
}
],
"runtime": {
"type": "docker",
"requiresGpu": true,
"dockerImage": "cr.internal.supervisely.com/ovalbee-internal/nodes/grounding_dino:0.0.17"
},
"widgets": {
"widget": {
"id": "asset-preview",
"settings": {
"datasetId": {
"type": "variable",
"value": "self.outputs.OUTPUT_DATASET"
}
}
}
}
}

References