Train Grounding DINO
Overview
Fine-tunes Grounding DINO on an OvalBee dataset. The node exports the datasets to COCO itself, so no conversion step runs in front of it.
Grounding DINO detects objects from text, which means the pretrained model already finds classes it was never trained on. Fine-tuning is what you do when that zero-shot behaviour is close but not accurate enough on your own data.
The badge shows the best mAP at IoU 0.5 reached on the validation split.
Pretrained Models
- Grounding DINO Swin-T (lighter, ~170M params)
- Grounding DINO Swin-B (~240M params)
- Grounding DINO Swin-L (best quality, ~340M params)
Prerequisites
What the node writes to disk
The node builds this tree in the task’s own scratch space at the start of every run, and data.yaml names the classes it trained on. Nothing here survives the run — the checkpoints leave through the model asset.
How it works
FAQ
Do I need a conversion step in front of this node?
No. Connect the datasets directly. The node exports them to COCO itself, in a single pass so the two splits agree on the class order.
Where do the trained files end up?
On the model asset. The run writes its checkpoints to scratch space that disappears with the container, so the best one is uploaded before the node finishes.
Is the model asset ready to use after training?
Yes. It carries the checkpoint, so predict and export nodes can take it as-is.
Can I watch the run while it trains?
Yes. The report is published early and keeps updating, so charts and checkpoints appear as they are produced rather than only at the end.
Do I need to train at all?
Often not. The pretrained model detects classes from their names, so try Predict Grounding DINO on your data first and fine-tune only if the zero-shot result falls short.
Which backbone should I pick?
Swin-T for speed and the smallest memory footprint, Swin-L when accuracy matters most, Swin-B in between. Larger backbones need a smaller batch size to fit.
Why did the run fail right at the start?
Usually the data or the config. Both datasets need annotated objects — a validation split with no boxes is refused rather than trained blind — and Config has to be a valid YAML mapping whose keys the trainer accepts.
Inputs
Dataset the model is trained on. The node converts it to COCO itself. Key: TRAIN_DATASET.
Dataset held out for validation. Grounding DINO reports mAP on it, and cannot train without it. Key: VAL_DATASET.
Model to fine-tune. A version copied from the zoo starts from the public weights of that architecture. One from an earlier run continues from its checkpoint. The run is registered as a new version of this model unless Destination Model names another one. Key: MODEL.
Where the trained version is registered, when it should not go to the model being trained. Leave empty to add the version to the Model above. Key: DESTINATION_MODEL.
Training hyperparameters in YAML - how long the run trains, at what resolution, and how much the images are augmented. Comments in the default explain the less obvious ones, and a commented-out line stays off until you uncomment it. Key: CONFIG.
Outputs
ID of the asset containing the live Grounding DINO training report. Shown as an artifact. Key: REPORT_ID.
Reference to the registry version this run produced, as <model id>@<version id>. It wires into the Model input of a predict or evaluate node and pins it to exactly these weights. Empty when the run could not be registered. Shown as an artifact. Key: MODEL.
Best mAP50 achieved during training (shown as a badge). Shown on the node as a badge. Key: BEST_MAP50.
Models and configuration
Model is both what training starts from and where its result is filed: the run adds a version to that model. A version copied from the zoo has no weights of its own, so the run starts from the public weights of that architecture. One produced by an earlier run continues from its checkpoint, and the new version records it as its parent.
Set Destination Model to file the result elsewhere — the trained version is created there instead, and Model is then only the starting weights.
The Swin-T backbone is the fastest, Swin-L the most accurate, and Swin-B the middle ground.
Config holds the training settings as YAML. Dataset paths and the class count are injected by the node.
Report
The report appears near the start of the run and updates as training produces results. It opens in the preview widget on the node.
Report sections
Runtime
The node runs in its own GPU container. Memory use follows the architecture, the input size, and the batch size, so a model that will not fit is the first thing to check when a run dies early.
The first run takes longer while the worker pulls the image and downloads pretrained weights. Every run downloads its datasets again — the container keeps nothing between runs.
JSON config
Machine-readable node interface for automation and advanced usage.