Skip to navigation

Label train, validation and test

Overview

Hands the images of a train, a validation and a test set to labelers from one queue, and keeps what they submit from each set apart: labeled train images, labeled validation images and labeled test images each land in a dataset of their own.

Each labeler works through the sets in order: train first, then validation, then test. Set any one, two or all three, and the node returns a labeled dataset for each one you set. Labelers see a single Labeled button and never pick where an image goes. Use it to label the splits a model trains and is scored on, where mixing them would let a model be scored on images it trained on. For custom status buttons or a single source, use Labeling instead.

Once every required input is set, the node runs itself every 15 minutes to top up queues and take back unfinished work.

Train/validation split Label train, validation and test Train

How it works

1

Checks each set you picked and the Classes ontology, and creates a labeled dataset for each set on its first run, reusing it on every later one.

2

Gives each person in Labelers a private queue, and sends the queue of anyone removed from the list back to the set it came from.

3

Tops each queue up to Batch size from the set it is on. A queue starts on the first set you picked and stays on a set until that set has no images left and the labeler has submitted everything they hold.

4

Moves an image and its annotation to its set’s labeled dataset as soon as the labeler submits it, then refills the queue right away, from the next set once this one runs out.

FAQ

The run fails when none of Train images, Validation images or Test images is set, when a set no longer exists, or when Classes names no ontology, since every queue is pinned to one.

A queue that is empty goes back to the first set, in order, that has images again. A queue still holding images of a later set finishes those first.

Images that sit in a queue longer than Return unclaimed work after (minutes) go back to the set they came from on the node’s next run, along with any annotation the labeler saved but never submitted.

On the next run, a queue on that set gives back what it holds and moves to the first set still picked. The labeled dataset of the cleared set keeps what was already submitted.

Every image still in a labeler’s queue goes back to the set it came from. The labeled datasets keep what was already submitted.

No. Each queue offers one Labeled button, and the vocabulary is the one picked in Classes. Use Labeling for either.

Inputs

Train images
dataset

Images to label for training. Labelers get these first. Anything left unsubmitted comes back here. Key: TRAIN_DATASET.

Validation images
dataset

Images to label as the validation set. Labelers get these once the train images run out. Key: VAL_DATASET.

Test images
dataset

Images to label as the test set. Labelers get these last. Key: TEST_DATASET.

Labelers
usersRequired

Workspace members who label. Each gets one private queue that works through the splits in order: train, then validation, then test. Key: LABELERS.

Classes
ontologyRequired

Vocabulary every split is labeled in. Defaults to the train images’ ontology. Tick only part of it to hide the rest from labelers, or tick nothing to label in all of it. Values come from TRAIN_DATASET. Key: CLASSES.

Batch size
integerDefaults to 10Required

Number of images a labeler holds at once. Their queue is topped up to this after each submit and on every node run. Minimum: 1. Key: BATCH_SIZE.

Return unclaimed work after (minutes)
integerDefaults to 120Required

Images left unsubmitted in a labeler’s queue longer than this go back to the split they came from, so nobody sits on work they never finished. Minimum: 1. Key: CLAIM_TTL_MINUTES.

Outputs

User queues
object

Internal map of user IDs to their hidden labeling queue datasets. Hidden from the flow editor. Key: USERS_DATASETS.

Outcome datasets
object

Internal list of the datasets submitted images land in. Hidden from the flow editor. Key: OUTCOME_DATASETS.

Labeled train
dataset

Dataset the labeled train images land in. Empty when no train images are set. Key: LABELED_TRAIN.

Labeled validation
dataset

Dataset the labeled validation images land in. Empty when no validation images are set. Key: LABELED_VAL.

Labeled test
dataset

Dataset the labeled test images land in. Empty when no test images are set. Key: LABELED_TEST.

Status
object

Compact per-labeler queue status table object. Hidden from the flow editor. Key: STATUS.

JSON config

Machine-readable node interface for automation and advanced usage.

{
"name": "Label train, validation and test",
"description": "Hand a train, a validation and a test set to labelers from one queue, split by split, and keep each split's labeled images in a dataset of its own.",
"category": "Label",
"namespace": "ovalbee",
"templateKey": "label/labeling_splits",
"version": "v1",
"inputs": [
{
"key": "TRAIN_DATASET",
"label": "Train images",
"type": "dataset",
"description": "Images to label for training. Labelers get these first. Anything left unsubmitted comes back here.",
"required": false,
"default": null,
"visibleWhen": null,
"options": {
"creatable": false
}
},
{
"key": "VAL_DATASET",
"label": "Validation images",
"type": "dataset",
"description": "Images to label as the validation set. Labelers get these once the train images run out.",
"required": false,
"default": null,
"visibleWhen": null,
"options": {
"creatable": false
}
},
{
"key": "TEST_DATASET",
"label": "Test images",
"type": "dataset",
"description": "Images to label as the test set. Labelers get these last.",
"required": false,
"default": null,
"visibleWhen": null,
"options": {
"creatable": false
}
},
{
"key": "LABELERS",
"label": "Labelers",
"type": "users",
"description": "Workspace members who label. Each gets one private queue that works through the splits in order: train, then validation, then test.",
"required": true,
"default": null,
"visibleWhen": null
},
{
"key": "CLASSES",
"label": "Classes",
"type": "ontology",
"description": "Vocabulary every split is labeled in. Defaults to the train images' ontology. Tick only part of it to hide the rest from labelers, or tick nothing to label in all of it.",
"required": true,
"default": null,
"visibleWhen": null,
"options": {
"ref": "TRAIN_DATASET",
"selectable": true,
"creatable": true
}
},
{
"key": "BATCH_SIZE",
"label": "Batch size",
"type": "number",
"description": "Number of images a labeler holds at once. Their queue is topped up to this after each submit and on every node run.",
"required": true,
"default": 10,
"visibleWhen": null,
"options": {
"type": "integer",
"min": 1
}
},
{
"key": "CLAIM_TTL_MINUTES",
"label": "Return unclaimed work after (minutes)",
"type": "number",
"description": "Images left unsubmitted in a labeler's queue longer than this go back to the split they came from, so nobody sits on work they never finished.",
"required": true,
"default": 120,
"visibleWhen": null,
"options": {
"type": "integer",
"min": 1
}
}
],
"outputs": [
{
"key": "USERS_DATASETS",
"label": "User queues",
"type": "object",
"description": "Internal map of user IDs to their hidden labeling queue datasets.",
"artifact": false,
"badge": false,
"badgeOpens": null,
"hidden": true
},
{
"key": "OUTCOME_DATASETS",
"label": "Outcome datasets",
"type": "object",
"description": "Internal list of the datasets submitted images land in.",
"artifact": false,
"badge": false,
"badgeOpens": null,
"hidden": true
},
{
"key": "LABELED_TRAIN",
"label": "Labeled train",
"type": "dataset",
"description": "Dataset the labeled train images land in. Empty when no train images are set.",
"artifact": false,
"badge": false,
"badgeOpens": null,
"hidden": false
},
{
"key": "LABELED_VAL",
"label": "Labeled validation",
"type": "dataset",
"description": "Dataset the labeled validation images land in. Empty when no validation images are set.",
"artifact": false,
"badge": false,
"badgeOpens": null,
"hidden": false
},
{
"key": "LABELED_TEST",
"label": "Labeled test",
"type": "dataset",
"description": "Dataset the labeled test images land in. Empty when no test images are set.",
"artifact": false,
"badge": false,
"badgeOpens": null,
"hidden": false
},
{
"key": "STATUS",
"label": "Status",
"type": "object",
"description": "Compact per-labeler queue status table object.",
"artifact": false,
"badge": false,
"badgeOpens": null,
"hidden": true
}
],
"automation": {
"requires_configuration": true,
"schedule": "*/15 * * * *",
"discard_history": "on_success"
},
"widgets": {
"widget": {
"id": "table",
"settings": {
"data": {
"type": "variable",
"value": "self.outputs.STATUS"
}
}
},
"secondaryWidget": {
"id": "open-labeling-tool",
"settings": {
"name": {
"type": "input",
"value": "Open annotation tool"
}
}
}
}
}