Train Grounding DINO

GPU Dockerfile
First run can take longer while the worker pulls the configured Docker image.

Overview

Fine-tunes Grounding DINO on an OvalBee dataset. The node exports the datasets to COCO itself, so no conversion step runs in front of it.

Grounding DINO detects objects from text, which means the pretrained model already finds classes it was never trained on. Fine-tuning is what you do when that zero-shot behaviour is close but not accurate enough on your own data.

The badge shows the best mAP at IoU 0.5 reached on the validation split.

Train dataset Train Grounding DINO Model asset Validation dataset
  • Grounding DINO Swin-T (lighter, ~170M params)
  • Grounding DINO Swin-B (~240M params)
  • Grounding DINO Swin-L (best quality, ~340M params)

Prerequisites

RequirementWhat you need
GPU workerA worker with a GPU.
Train datasetAn annotated dataset with the objects to learn.
Validation datasetA second annotated dataset. Grounding DINO scores mAP on it every val_interval epochs and refuses to train without it.
A model to fine-tuneA model of this workspace, picked in Model. Create one from the picker to copy an architecture in from the zoo.
📁 dataset_directory/
├── 📄 data.yaml
├── 📁 images/
│ ├── 📁 train/
│ └── 📁 val/
└── 📁 annotations/
├── 📄 train.json
└── 📄 val.json

The node builds this tree in the task’s own scratch space at the start of every run, and data.yaml names the classes it trained on. Nothing here survives the run — the checkpoints leave through the model asset.

How it works

1

Exports both datasets to COCO in one pass, so train and val share one class index.

2

Builds the training config from the selected architecture, the class list, and your overrides.

3

Fine-tunes Grounding DINO, writing checkpoints, metrics, and TensorBoard events beside the dataset.

4

Uploads the best checkpoint and creates the model asset that carries it, along with the architecture, classes, and final metrics.

FAQ

No. Connect the datasets directly. The node exports them to COCO itself, in a single pass so the two splits agree on the class order.

On the model asset. The run writes its checkpoints to scratch space that disappears with the container, so the best one is uploaded before the node finishes.

Yes. It carries the checkpoint, so predict and export nodes can take it as-is.

Yes. The report is published early and keeps updating, so charts and checkpoints appear as they are produced rather than only at the end.

Often not. The pretrained model detects classes from their names, so try Predict Grounding DINO on your data first and fine-tune only if the zero-shot result falls short.

Swin-T for speed and the smallest memory footprint, Swin-L when accuracy matters most, Swin-B in between. Larger backbones need a smaller batch size to fit.

Usually the data or the config. Both datasets need annotated objects — a validation split with no boxes is refused rather than trained blind — and Config has to be a valid YAML mapping whose keys the trainer accepts.

Inputs

Train dataset
datasetRequired

Dataset the model is trained on. The node converts it to COCO itself. Key: TRAIN_DATASET.

Validation dataset
datasetRequired

Dataset held out for validation. Grounding DINO reports mAP on it, and cannot train without it. Key: VAL_DATASET.

Model
modelRequired

Model to fine-tune. A version copied from the zoo starts from the public weights of that architecture. One from an earlier run continues from its checkpoint. The run is registered as a new version of this model unless Destination Model names another one. Key: MODEL.

Destination model
model

Where the trained version is registered, when it should not go to the model being trained. Leave empty to add the version to the Model above. Key: DESTINATION_MODEL.

Hyperparameters
yamlRequired

Training hyperparameters in YAML - how long the run trains, at what resolution, and how much the images are augmented. Comments in the default explain the less obvious ones, and a commented-out line stays off until you uncomment it. Key: CONFIG.

Outputs

Report Asset
report

ID of the asset containing the live Grounding DINO training report. Shown as an artifact. Key: REPORT_ID.

Model
model

Reference to the registry version this run produced, as <model id>@<version id>. It wires into the Model input of a predict or evaluate node and pins it to exactly these weights. Empty when the run could not be registered. Shown as an artifact. Key: MODEL.

Best mAP50
string

Best mAP50 achieved during training (shown as a badge). Shown on the node as a badge. Key: BEST_MAP50.

Models and configuration

Model is both what training starts from and where its result is filed: the run adds a version to that model. A version copied from the zoo has no weights of its own, so the run starts from the public weights of that architecture. One produced by an earlier run continues from its checkpoint, and the new version records it as its parent.

Set Destination Model to file the result elsewhere — the trained version is created there instead, and Model is then only the starting weights.

The Swin-T backbone is the fastest, Swin-L the most accurate, and Swin-B the middle ground.

Config holds the training settings as YAML. Dataset paths and the class count are injected by the node.

Report

The report appears near the start of the run and updates as training produces results. It opens in the preview widget on the node.

SectionWhat it shows
OverviewThe architecture, task, configured epochs, elapsed time, dataset sizes, and classes.
PredictionsThe best checkpoint applied to a fixed sample of validation images, once one exists.
Evaluation MetricsDetection metrics from the best validation epoch.
Training PlotsEvery loss, metric, and learning-rate series recorded during the run.
TensorBoardThe raw TensorBoard event stream.
ArtifactsThe checkpoints written so far, and the model asset panel once the run finishes.
ClassesThe class names the run trained on.
HyperparametersThe config the run actually used.
How to Use the Trained ModelA short local example that loads the checkpoint and predicts an image.

Runtime

The node runs in its own GPU container. Memory use follows the architecture, the input size, and the batch size, so a model that will not fit is the first thing to check when a run dies early.

The first run takes longer while the worker pulls the image and downloads pretrained weights. Every run downloads its datasets again — the container keeps nothing between runs.

JSON config

Machine-readable node interface for automation and advanced usage.

{
"name": "Train Grounding DINO",
"description": "Train Grounding DINO object detection model using MMDetection framework.",
"category": "Train",
"namespace": "ovalbee",
"templateKey": "models/grounding_dino/train_grounding_dino",
"version": "v1",
"inputs": [
{
"key": "TRAIN_DATASET",
"label": "Train dataset",
"type": "dataset",
"description": "Dataset the model is trained on. The node converts it to COCO itself.",
"required": true,
"default": null,
"visibleWhen": null,
"options": {
"creatable": false
}
},
{
"key": "VAL_DATASET",
"label": "Validation dataset",
"type": "dataset",
"description": "Dataset held out for validation. Grounding DINO reports mAP on it, and cannot train without it.",
"required": true,
"default": null,
"visibleWhen": null,
"options": {
"creatable": false
}
},
{
"key": "MODEL",
"label": "Model",
"type": "model",
"description": "Model to fine-tune. A version copied from the zoo starts from the public weights of that architecture. One from an earlier run continues from its checkpoint. The run is registered as a new version of this model unless Destination Model names another one.\n",
"required": true,
"default": null,
"visibleWhen": null,
"options": {
"scope": "my",
"framework": "grounding_dino",
"task_type": "object_detection"
}
},
{
"key": "DESTINATION_MODEL",
"label": "Destination model",
"type": "model",
"description": "Where the trained version is registered, when it should not go to the model being trained. Leave empty to add the version to the Model above.\n",
"required": false,
"default": null,
"visibleWhen": null,
"options": {
"scope": "my",
"framework": "grounding_dino",
"task_type": "object_detection"
}
},
{
"key": "CONFIG",
"label": "Hyperparameters",
"type": "yaml",
"description": "Training hyperparameters in YAML - how long the run trains, at what resolution, and how much the images are augmented. Comments in the default explain the less obvious ones, and a commented-out line stays off until you uncomment it.\n",
"required": true,
"default": "# --- Schedule ---\nepochs: 12\nbatch_size: 2\nlr: 5.0e-5\nweight_decay: 0.0001\nwarmup_iters: 200\nval_interval: 1\ncheckpoint_interval: 5\nworkers: 4\n# --- Input ---\nimgsz: 800 # input resolution: scales whole multi-scale envelope by imgsz/800 (aspect kept). short edge up to imgsz, long-edge cap ~imgsz*1.667. For 16:9 the long cap binds, so e.g. native 1920x1080 ~ imgsz 1150.\n# --- Augmentation ---\nrandom_flip: 0.5 # horizontal flip probability (0.0 to disable)\n# random_flip_vertical: 0.0 # vertical flip probability (0.0 to disable)\n# scale_min: 0.5 # min random resize scale factor relative to imgsz\n# scale_max: 2.0 # max random resize scale factor relative to imgsz\n# brightness_delta: 32 # max pixel brightness shift (0-255, 0 to disable)\n# contrast_lower: 0.5 # contrast multiplier lower bound\n# contrast_upper: 1.5 # contrast multiplier upper bound\n# saturation_lower: 0.5 # saturation multiplier lower bound\n# saturation_upper: 1.5 # saturation multiplier upper bound\n# hue_delta: 18 # max hue shift in degrees (0 to disable)\n# --- Model ---\nfreeze_backbone: false # freeze image backbone weights\nfreeze_text_encoder: false # freeze BERT text encoder; set true only to save VRAM (reduces quality for overlapping class names)\n# --- LR Schedule ---\n# lr_schedule: cosine # \"constant\" (default) or \"cosine\" - cosine decay from lr to lr*lr_min_ratio\n# lr_min_ratio: 0.01 # final LR as a fraction of lr (only used with lr_schedule: cosine)\n# --- Optimizer ---\ngrad_clip: 0.1 # gradient clip max norm (0 to disable)\n# --- Performance ---\namp: false # mixed precision (AmpOptimWrapper): ~half activation memory + faster on tensor-core GPUs, fits a larger batch. Set true to enable.\n# amp_dtype: fp16 # \"fp16\" (default, required: mmcv deformable-attn has no bf16 CUDA kernel) or \"bf16\" (only if a future mmcv build adds the bf16 MSDA kernel)\n",
"visibleWhen": null
}
],
"outputs": [
{
"key": "REPORT_ID",
"label": "Report Asset",
"type": "report",
"description": "ID of the asset containing the live Grounding DINO training report.",
"artifact": true,
"badge": false,
"badgeOpens": null,
"hidden": false
},
{
"key": "MODEL",
"label": "Model",
"type": "model",
"description": "Reference to the registry version this run produced, as `<model id>@<version id>`. It wires into the Model input of a predict or evaluate node and pins it to exactly these weights. Empty when the run could not be registered.\n",
"artifact": true,
"badge": false,
"badgeOpens": null,
"hidden": false
},
{
"key": "BEST_MAP50",
"label": "Best mAP50",
"type": "string",
"description": "Best mAP50 achieved during training (shown as a badge).",
"artifact": false,
"badge": true,
"badgeOpens": null,
"hidden": false
}
],
"runtime": {
"type": "docker",
"requiresGpu": true,
"dockerImage": "cr.internal.supervisely.com/ovalbee-internal/nodes/grounding_dino:0.0.16"
},
"widgets": {
"widget": {
"id": "report-preview",
"settings": {
"reportAsset": {
"type": "variable",
"value": "self.outputs.REPORT_ID"
},
"name": {
"type": "input",
"value": "Open training report"
}
}
}
}
}

References