Serve live training with Grounding DINO
Overview
Trains a Grounding DINO detector continuously while people are still labeling, and predicts with it at the same time. Submit an image the moment it is labeled and the detector learns from it within a couple of iterations, so the boxes it suggests on the next image come from a model that has already seen the corrections made on the last one.
This is a deployment, not a flow node. It is started from the Deployments page, holds one GPU for its whole lifetime, and is driven over HTTP — by the labeling tool as images are submitted, or by anything else that can call the routes below. Reach for it when the wait between a label and its effect on the suggestions is what you are trying to remove. For a fine-tune with a fixed budget that ends, use Train Grounding DINO.
Prerequisites
How it works
Loads the weights the Model input names, or the public checkpoint of the architecture it names when the version carries no weights of its own.
Starts an endless fine-tune on a background thread. It waits there until initial_samples
labeled images have arrived.
Takes labeled images on POST /samples. Each one is downloaded, its newest annotation is read,
and its boxes join the training set. A class nobody had drawn yet is appended to the vocabulary.
FAQ
How do I call the deployment?
Every endpoint is reached through the deployment’s own invoke URL, which the Deployments page
shows on the deployment’s info dialog. GET /status reports the phase, the number of images
submitted, the iteration and loss, and the vocabulary the session is training in.
What happens when the model has learned all it can from the current images?
Training pauses. The plateau_* keys watch the loss over two moving windows, and once it stops
falling the session stops burning GPU on data it has exhausted. The next submitted image
resumes it. Set plateau_window: 0 to train without ever pausing.
Can it predict before anything has been labeled?
Only if classes in Config lists the class names up front. Grounding DINO is prompted with
class names, so a session seeded that way pre-labels from the very first image using the public
weights. Left empty, the vocabulary is learned from submissions and POST /predict answers 400
until the first labeled image arrives.
What happens if I submit the same image twice?
Its boxes are replaced by the newer ones. That is what makes correcting an image worth doing: the session trains on the corrected version, not on both.
Where do the trained weights go?
Into the Model input’s own family in the Model Registry, as one version that is advanced on every publish rather than a new version each time. It is ready from the first publish, so a predict or train node downstream can pick up the session’s weights while it is still running.
Does stopping the deployment lose the training?
Everything trained since the last publish. A stopping container is given ten seconds, which is
not enough to write and upload a detector’s weights, so the publish the deployment attempts on
its way down usually does not finish. publish_every_minutes is what actually bounds the loss.
Call POST /checkpoint before stopping a session you care about.
What does this node not support?
One GPU and one session per deployment. It trains bounding boxes only, so a polygon or mask label is learned as the box around it. It holds no validation split and reports no metrics beyond the training loss, because nothing is held out. The submitted images live in the container, so a restarted deployment starts from the last registered version with an empty training set.
Inputs
Model the session starts from, and the model each checkpoint is registered as a new version of. Copy a Grounding DINO entry from the zoo into this workspace to get one, a version with no weights of its own starts from that architecture’s public checkpoint. Key: MODEL.
Where the fine-tune runs. CPU is only usable for a smoke test - one training iteration takes minutes there. Key: INFERENCE_DEVICE.
Options:
- GPU (CUDA) (
cuda) - CPU (
cpu)
YAML settings for the background fine-tune and for prediction. The vocabulary comes from the ontology of the dataset the samples are submitted from, so the session can pre-label every class from the first image without anyone listing them. classes adds names that ontology does not carry, and every class drawn on a submitted image is appended too. The plateau_* keys pause training once the loss stops falling, until a new sample arrives. publish_every_minutes is how often the running weights are registered as a version. Key: CONFIG.
Outputs
Live report for the session: loss curves, how close its pre-labels are getting to what the labeler submits, and a timeline of every submission, plateau and checkpoint. It keeps updating for as long as the session is up. Shown as an artifact. Key: REPORT_ID.
The weights the session has trained so far, as <model id>@<version id>. Wire it into the Model input of a predict or evaluate node to label the rest of the dataset once the session is good enough. It appears after the first checkpoint and keeps pointing at the latest weights, so a run started later uses what the session knows by then. Shown as an artifact. Key: MODEL.
The running session. Wire it into a node that takes a deployment to send it samples or prediction requests while it is up. Key: DEPLOYMENT.
Models and configuration
The Model input is resolved to one version before the container starts. A version with its own checkpoint is loaded from that file. A version copied from the zoo with no weights of its own resolves to the public MM Grounding DINO checkpoint for its architecture, one of Swin-T, Swin-B or Swin-L, which is downloaded on first use.
Config carries the whole fine-tune. initial_samples is how many labeled images must arrive before
training starts. batch_size, lr, weight_decay, grad_clip, warmup_iters and amp set the
optimizer, and imgsz, random_flip, freeze_backbone and freeze_text_encoder mirror the Train
Grounding DINO node. imgsz applies to both training and prediction, so the boxes the annotator
reviews come from the model at the resolution it was trained at. conf_threshold is the lowest
score a prediction keeps, and a POST /predict call may override that one key for itself.
Runtime
Runs as a Docker deployment on the grounding_dino node image, which carries MMDetection 3.3 and
the MM Grounding DINO configs, and requests a GPU. The fine-tune runs on a thread inside the same
process that answers requests, so the dataset it trains on is held in memory and the container
never reads it back from disk between iterations.
JSON config
Machine-readable node interface for automation and advanced usage.