Serve live training with Grounding DINO

GPU Dockerfile
First run can take longer while the worker pulls the configured Docker image.

Overview

Trains a Grounding DINO detector continuously while people are still labeling, and predicts with it at the same time. Submit an image the moment it is labeled and the detector learns from it within a couple of iterations, so the boxes it suggests on the next image come from a model that has already seen the corrections made on the last one.

This is a deployment, not a flow node. It is started from the Deployments page, holds one GPU for its whole lifetime, and is driven over HTTP — by the labeling tool as images are submitted, or by anything else that can call the routes below. Reach for it when the wait between a label and its effect on the suggestions is what you are trying to remove. For a fine-tune with a fixed budget that ends, use Train Grounding DINO.

Annotator labels an image POST /samples Background fine-tune POST /predict Model version

Prerequisites

RequirementWhat you need
GPU workerThe fine-tune and the predictions share one GPU for as long as the deployment runs.
Grounding DINO model in this workspaceThe Model input only offers this workspace’s own models. Copy a Grounding DINO entry from the zoo to get one.
A client that calls the deploymentNothing in the flow editor drives this node. Its endpoints are called over HTTP, from a script or from a labeling integration.
Labeled images in a datasetEvery image submitted for training is read back from the workspace, so its annotation has to be saved before the id is sent.

How it works

1

Loads the weights the Model input names, or the public checkpoint of the architecture it names when the version carries no weights of its own.

2

Starts an endless fine-tune on a background thread. It waits there until initial_samples labeled images have arrived.

3

Takes labeled images on POST /samples. Each one is downloaded, its newest annotation is read, and its boxes join the training set. A class nobody had drawn yet is appended to the vocabulary.

4

Answers POST /predict with the weights as they stand, pausing training for the length of one forward pass so a prediction never sees a half-applied update.

5

Registers the weights as a version of the Model every publish_every_minutes and on POST /checkpoint, advancing one version rather than adding one per publish.

FAQ

Every endpoint is reached through the deployment’s own invoke URL, which the Deployments page shows on the deployment’s info dialog. GET /status reports the phase, the number of images submitted, the iteration and loss, and the vocabulary the session is training in.

Training pauses. The plateau_* keys watch the loss over two moving windows, and once it stops falling the session stops burning GPU on data it has exhausted. The next submitted image resumes it. Set plateau_window: 0 to train without ever pausing.

Only if classes in Config lists the class names up front. Grounding DINO is prompted with class names, so a session seeded that way pre-labels from the very first image using the public weights. Left empty, the vocabulary is learned from submissions and POST /predict answers 400 until the first labeled image arrives.

Its boxes are replaced by the newer ones. That is what makes correcting an image worth doing: the session trains on the corrected version, not on both.

Into the Model input’s own family in the Model Registry, as one version that is advanced on every publish rather than a new version each time. It is ready from the first publish, so a predict or train node downstream can pick up the session’s weights while it is still running.

Everything trained since the last publish. A stopping container is given ten seconds, which is not enough to write and upload a detector’s weights, so the publish the deployment attempts on its way down usually does not finish. publish_every_minutes is what actually bounds the loss. Call POST /checkpoint before stopping a session you care about.

One GPU and one session per deployment. It trains bounding boxes only, so a polygon or mask label is learned as the box around it. It holds no validation split and reports no metrics beyond the training loss, because nothing is held out. The submitted images live in the container, so a restarted deployment starts from the last registered version with an empty training set.

Inputs

Model
modelRequired

Model the session starts from, and the model each checkpoint is registered as a new version of. Copy a Grounding DINO entry from the zoo into this workspace to get one, a version with no weights of its own starts from that architecture’s public checkpoint. Key: MODEL.

Device
selectDefaults to cudaRequired

Where the fine-tune runs. CPU is only usable for a smoke test - one training iteration takes minutes there. Key: INFERENCE_DEVICE.

Options:

  • GPU (CUDA) (cuda)
  • CPU (cpu)
Config
yamlRequired

YAML settings for the background fine-tune and for prediction. The vocabulary comes from the ontology of the dataset the samples are submitted from, so the session can pre-label every class from the first image without anyone listing them. classes adds names that ontology does not carry, and every class drawn on a submitted image is appended too. The plateau_* keys pause training once the loss stops falling, until a new sample arrives. publish_every_minutes is how often the running weights are registered as a version. Key: CONFIG.

Outputs

Report
report

Live report for the session: loss curves, how close its pre-labels are getting to what the labeler submits, and a timeline of every submission, plateau and checkpoint. It keeps updating for as long as the session is up. Shown as an artifact. Key: REPORT_ID.

Model
model

The weights the session has trained so far, as <model id>@<version id>. Wire it into the Model input of a predict or evaluate node to label the rest of the dataset once the session is good enough. It appears after the first checkpoint and keeps pointing at the latest weights, so a run started later uses what the session knows by then. Shown as an artifact. Key: MODEL.

Deployment
deployment

The running session. Wire it into a node that takes a deployment to send it samples or prediction requests while it is up. Key: DEPLOYMENT.

Models and configuration

The Model input is resolved to one version before the container starts. A version with its own checkpoint is loaded from that file. A version copied from the zoo with no weights of its own resolves to the public MM Grounding DINO checkpoint for its architecture, one of Swin-T, Swin-B or Swin-L, which is downloaded on first use.

Config carries the whole fine-tune. initial_samples is how many labeled images must arrive before training starts. batch_size, lr, weight_decay, grad_clip, warmup_iters and amp set the optimizer, and imgsz, random_flip, freeze_backbone and freeze_text_encoder mirror the Train Grounding DINO node. imgsz applies to both training and prediction, so the boxes the annotator reviews come from the model at the resolution it was trained at. conf_threshold is the lowest score a prediction keeps, and a POST /predict call may override that one key for itself.

Runtime

Runs as a Docker deployment on the grounding_dino node image, which carries MMDetection 3.3 and the MM Grounding DINO configs, and requests a GPU. The fine-tune runs on a thread inside the same process that answers requests, so the dataset it trains on is held in memory and the container never reads it back from disk between iterations.

JSON config

Machine-readable node interface for automation and advanced usage.

{
"name": "Serve live training with Grounding DINO",
"description": "Keeps a Grounding DINO fine-tune running for as long as the deployment is up, learning from labeled images as they are submitted and answering prediction requests from the model being trained. Holds the GPU for its whole lifetime.",
"category": "Models",
"namespace": "ovalbee",
"templateKey": "models/grounding_dino/serve_live_training_with_grounding_dino",
"version": "v1",
"inputs": [
{
"key": "MODEL",
"label": "Model",
"type": "model",
"description": "Model the session starts from, and the model each checkpoint is registered as a new version of. Copy a Grounding DINO entry from the zoo into this workspace to get one; a version with no weights of its own starts from that architecture's public checkpoint.\n",
"required": true,
"default": null,
"visibleWhen": null,
"options": {
"scope": "my",
"framework": "grounding_dino",
"task_type": "object_detection"
}
},
{
"key": "INFERENCE_DEVICE",
"label": "Device",
"type": "select",
"description": "Where the fine-tune runs. CPU is only usable for a smoke test - one training iteration takes minutes there.\n",
"required": true,
"default": "cuda",
"visibleWhen": null,
"options": {
"options": [
{
"label": "GPU (CUDA)",
"value": "cuda"
},
{
"label": "CPU",
"value": "cpu"
}
]
}
},
{
"key": "CONFIG",
"label": "Config",
"type": "yaml",
"description": "YAML settings for the background fine-tune and for prediction. The vocabulary comes from the ontology of the dataset the samples are submitted from, so the session can pre-label every class from the first image without anyone listing them. classes adds names that ontology does not carry, and every class drawn on a submitted image is appended too. The plateau_* keys pause training once the loss stops falling, until a new sample arrives. publish_every_minutes is how often the running weights are registered as a version.\n",
"required": true,
"default": "# --- Vocabulary ---\nclasses: [] # extra classes to prompt with; the dataset's ontology is used automatically\n# --- Schedule ---\ninitial_samples: 2 # submitted images required before training starts\nwarmup_iters: 100\nbatch_size: 2\nlr: 5.0e-5\nweight_decay: 0.0001\ngrad_clip: 0.1 # gradient clip max norm (0 to disable)\namp: false # fp16 mixed precision: less VRAM, faster on tensor-core GPUs\n# --- Pause when the loss stalls ---\nplateau_window: 150 # iterations per moving-average window (0 to never pause)\nplateau_threshold: 0.002 # smallest mean-loss drop between windows that still counts as progress\nplateau_patience: 3 # consecutive stalled checks before pausing\n# --- Input ---\nimgsz: 800 # input resolution (short-edge reference, long edge ~imgsz*1.667)\nrandom_flip: 0.5 # horizontal flip probability (0.0 to disable)\n# --- Model ---\nfreeze_backbone: false # freeze the image backbone (faster, less VRAM)\nfreeze_text_encoder: false # freeze BERT. Keep false so class names with shared tokens separate\n# --- Prediction ---\nconf_threshold: 0.3\n# --- Checkpoints ---\npublish_every_minutes: 10 # 0 to publish only when the deployment stops\n",
"visibleWhen": null
}
],
"outputs": [
{
"key": "REPORT_ID",
"label": "Report",
"type": "report",
"description": "Live report for the session: loss curves, how close its pre-labels are getting to what the labeler submits, and a timeline of every submission, plateau and checkpoint. It keeps updating for as long as the session is up.\n",
"artifact": true,
"badge": false,
"badgeOpens": null,
"hidden": false
},
{
"key": "MODEL",
"label": "Model",
"type": "model",
"description": "The weights the session has trained so far, as `<model id>@<version id>`. Wire it into the Model input of a predict or evaluate node to label the rest of the dataset once the session is good enough. It appears after the first checkpoint and keeps pointing at the latest weights, so a run started later uses what the session knows by then.\n",
"artifact": true,
"badge": false,
"badgeOpens": null,
"hidden": false
},
{
"key": "DEPLOYMENT",
"label": "Deployment",
"type": "deployment",
"description": "The running session. Wire it into a node that takes a deployment to send it samples or prediction requests while it is up.\n",
"artifact": false,
"badge": false,
"badgeOpens": null,
"hidden": false
}
],
"runtime": {
"type": "docker",
"requiresGpu": true,
"dockerImage": "cr.internal.supervisely.com/ovalbee-internal/nodes/grounding_dino:0.0.16"
},
"widgets": {
"widget": {
"id": "report-preview",
"settings": {
"reportAsset": {
"type": "variable",
"value": "self.outputs.REPORT_ID"
}
}
}
}
}

References