Prelabel with live training

Overview

Pre-labels a whole dataset with a live training session that is still running, so labelers open images that already carry boxes from the latest weights.

Use it to pre-label a batch before it goes to labeling: the session keeps learning from every submitted image, and each batch it pre-labels comes from a model that has seen everything labeled so far. Unlike a Predict node, nothing has to stop the session to produce a checkpoint first.

Serve live training Prelabel with live training Labeling

Prerequisites

RequirementWhat you need
Running live training sessionA session started from the Serve live training with Grounding DINO node or the Deployments page. It must already know its classes, from a submitted image or from its Config.
Images in the workspaceA dataset of images. Videos are not pre-labeled.

How it works

1

Creates a new dataset on the input dataset’s classes and copies the images into it.

2

Sends the images to the session in batches. Each batch is predicted between two of the session’s training steps, with the weights it has at that moment.

3

Saves each image’s predicted boxes as its annotation in the new dataset and reports how many images were pre-labeled.

FAQ

It does not stop it. Training pauses for each request while the batch is predicted and resumes right after. A smaller Images per request hands the GPU back to training sooner.

The session refuses to predict while it is still loading, or before it knows any classes. Wait until it is up, or submit a labeled image or list classes in its Config, and run the node again.

Yes. The session remembers every prediction it hands out, so when a labeler submits a pre-labeled image, the session scores the submission against the boxes this node saved.

Nothing. The images are copied into the new dataset with the predictions as their annotations. Connect the new dataset to the step that queues images for labeling.

Inputs

Live training session
deploymentRequired

Running live training session to ask. It keeps training while this node runs, and each image is predicted with the weights it has at that moment. Key: DEPLOYMENT.

Dataset
datasetRequired

Images to pre-label. They are left untouched here and copied, with the predicted boxes, into a new dataset. Key: DATASET.

Confidence threshold
number

Keep only boxes at least this confident. Empty uses the session’s own threshold. Minimum: 0. Maximum: 1. Step: 0.05. Key: CONF_THRESHOLD.

Images per request
integerDefaults to 8Required

How many images each request to the session carries. The session pauses training for every request, so a smaller batch hands the GPU back to training sooner. Minimum: 1. Maximum: 64. Step: 1. Key: BATCH_SIZE.

Outputs

Prelabeled dataset
dataset

New dataset holding the input images with the session’s predictions as their annotations. Key: OUTPUT_DATASET.

Prelabeled images
string

How many images got predictions, and the training iteration of the weights that made the last of them. Shown on the node as a badge. Selecting the badge opens OUTPUT_DATASET. Key: PRED_COUNT.

Models and configuration

The model is whatever the session is training at the moment of each request, so there is no model to pick here. The session’s own Config decides the vocabulary. Confidence threshold overrides the session’s threshold for this run only, and leaving it empty keeps the session’s.

Runtime

Runs on an ordinary worker without a GPU. The predictions run on the session’s GPU, so the session has to be up for the whole run.

JSON config

Machine-readable node interface for automation and advanced usage.

{
"name": "Prelabel with live training",
"description": "Pre-label a dataset's images with the weights a running live training session has right now, into a new dataset, without stopping the session.",
"category": "Predict",
"namespace": null,
"templateKey": "models/general/prelabel_with_live_training",
"version": "v1",
"inputs": [
{
"key": "DEPLOYMENT",
"label": "Live training session",
"type": "deployment",
"description": "Running live training session to ask. It keeps training while this node runs, and each image is predicted with the weights it has at that moment.",
"required": true,
"default": null,
"visibleWhen": null,
"options": {
"kind": "live_train"
}
},
{
"key": "DATASET",
"label": "Dataset",
"type": "dataset",
"description": "Images to pre-label. They are left untouched here and copied, with the predicted boxes, into a new dataset.",
"required": true,
"default": null,
"visibleWhen": null,
"options": {
"creatable": false
}
},
{
"key": "CONF_THRESHOLD",
"label": "Confidence threshold",
"type": "number",
"description": "Keep only boxes at least this confident. Empty uses the session's own threshold.",
"required": false,
"default": null,
"visibleWhen": null,
"options": {
"min": 0,
"max": 1,
"step": 0.05,
"type": "float"
}
},
{
"key": "BATCH_SIZE",
"label": "Images per request",
"type": "number",
"description": "How many images each request to the session carries. The session pauses training for every request, so a smaller batch hands the GPU back to training sooner.",
"required": true,
"default": 8,
"visibleWhen": null,
"options": {
"min": 1,
"max": 64,
"step": 1,
"type": "integer"
}
}
],
"outputs": [
{
"key": "OUTPUT_DATASET",
"label": "Prelabeled dataset",
"type": "dataset",
"description": "New dataset holding the input images with the session's predictions as their annotations.",
"artifact": false,
"badge": false,
"badgeOpens": null,
"hidden": false
},
{
"key": "PRED_COUNT",
"label": "Prelabeled images",
"type": "string",
"description": "How many images got predictions, and the training iteration of the weights that made the last of them.",
"artifact": false,
"badge": true,
"badgeOpens": "OUTPUT_DATASET",
"hidden": false
}
],
"automation": {
"requires_configuration": true
},
"widgets": {
"widget": {
"id": "asset-preview",
"settings": {
"datasetId": {
"type": "variable",
"value": "self.outputs.OUTPUT_DATASET"
}
}
}
}
}

References