Predict Grounding DINO
Overview
Runs Grounding DINO over a dataset and saves the detections as annotations in the output dataset.
Pick a Destination dataset to add the predictions to a dataset you already have, such as a labeling pool or an active-learning set: it keeps what it holds, and an image it already annotates gets the new prediction instead. Leave it empty and the node writes into a dataset of its own, one version per run.
The model detects objects from their names, so a public checkpoint finds your classes without any training. That makes it the fastest way to pre-label a fresh dataset.
The badge counts the annotated images and opens the output dataset.
Pretrained Models
- Grounding DINO Swin-T (lighter, ~170M params)
- Grounding DINO Swin-B (~240M params)
- Grounding DINO Swin-L (best quality, ~340M params)
Prerequisites
How it works
FAQ
Does it change the input dataset?
No. Predictions land in the output dataset and the source keeps whatever it already had.
What does Merge mode decide?
Replace leaves only the new predictions in the output dataset, which is what you want before an evaluation node. Keep both carries the input annotations over next to the predictions, which is what you want when comparing them side by side.
Can it detect classes nobody trained it on?
Yes, that is the point of the model. It matches objects against the class names it is given, so a public checkpoint already handles most everyday objects. Accuracy drops on domain-specific or visually unusual classes, which is when fine-tuning pays off.
How do the class names affect the result?
They are what the model matches against, so wording matters. A plain, concrete noun works better than an internal code or an abbreviation.
When should I use a text prompt instead of a class vocabulary?
A vocabulary is the usual choice: the classes are reusable, the output dataset is pinned to them, and a later run of the same flow detects the same things. A text prompt suits a one-off look at a dataset, or a description no class name captures — “a person wearing a yellow helmet”. Only the noun phrases of the sentence can be predicted, so anything it does not name is not detected. Typing short names separated by commas is treated as a plain class list, not as a sentence.
How do I know whether the predictions are any good?
Run the matching evaluation node against a ground-truth dataset. Prediction alone reports how many images were annotated, not how correct the annotations are.
Can I run this without choosing the framework by hand?
Yes, through Apply model. It reads the framework off the model asset and dispatches to this node, which is what you want when the model is only picked at runtime.
Inputs
Dataset to run detection on. Images are annotated when the dataset holds any, videos frame by frame otherwise. Key: DATASET.
Dataset the predictions are added to, keeping what it already holds. An image it already annotates gets the new prediction instead. Leave empty to use the node’s own output dataset, which a rerun writes a new version of. Key: DST_DATASET.
Model to run inference with, from this workspace or the zoo. The node is handed the resolved version and runs on its weights: the checkpoint stored for it, or the public weights of its architecture for a pretrained entry that hosts none. Key: MODEL.
Grounding DINO is prompted with text, so what it looks for is chosen here rather than baked into the weights. A vocabulary asks for a list of classes. A text prompt grounds the noun phrases of a sentence. Key: PROMPT_MODE.
Options:
- Class vocabulary (
vocabulary) - Text prompt (
prompt)
Vocabulary to detect - every class is asked for, and ticking a subset asks for that subset only. Tags are never asked for, since a tag has no box. Type a name to add a class the ontology does not have yet, or leave it empty to detect what a fine-tuned checkpoint was trained on. Visible when PROMPT_MODE is vocabulary. Key: CLASSES.
Sentence to ground, for example - The picture contains watermelon, flower, and a white bottle. Its noun phrases become the predicted classes, so a plain comma-separated list such as cat, dog, car is read as those three classes instead. Placeholder: There is a person wearing a yellow helmet.. Visible when PROMPT_MODE is prompt. Key: PROMPT.
Inference settings in YAML - how confident a detection must be to be kept, how many images go through the GPU at once, and the resolution they are run at. Set that resolution to the one the model was trained at, so training and inference see images the same way. Key: CONFIG.
Number of images fetched at a time while the GPU works through the ones already downloaded. Minimum: 1. Key: DOWNLOAD_BATCH_SIZE.
How far downloading may run ahead of inference, counted in batches. Raise it when a slow connection leaves the GPU idle, at the cost of disk space. Minimum: 1. Key: PREFETCH_BATCHES.
Outputs
Dataset with the input images plus the Grounding DINO prediction annotations: the destination dataset when one is picked, otherwise the node’s own. Shown as an artifact. Key: OUTPUT_DATASET.
Number of assets annotated. Shown on the node as a badge. Selecting the badge opens OUTPUT_DATASET. Key: PRED_COUNT.
Models and configuration
Model is the one place the weights come from: a model of this workspace, or one of the zoo entries listed above. The node is handed the resolved version and runs on its weights — the checkpoint stored for it, or the public weights of its architecture for an entry that hosts none.
What to detect decides where the prompt comes from, because the weights hold no class list of their own. With Class vocabulary the classes of the picked ontology are the prompt, and ticking a subset asks for that subset only. Leaving it empty falls back to the classes a fine-tuned checkpoint was trained on. With Text prompt the sentence is the prompt and its noun phrases become the predicted classes.
The output dataset is pinned to the vocabulary the predictions are named in, so the classes stay readable in the labeling tool and in evaluation.
Config holds the inference settings. conf_threshold is the score a detection needs to survive, and batch_size how many images the model sees at once.
Runtime
The node runs in its own GPU container. The first run takes longer while the worker pulls the image and downloads the selected weights.
Images are fetched ahead of the model to keep the GPU busy. Download Batch Size sets how many arrive at once and Prefetch Batches how many wait on local disk, so raise them for fast storage and lower them when disk is tight.
JSON config
Machine-readable node interface for automation and advanced usage.