Label train, validation and test
Overview
Hands the images of a train, a validation and a test set to labelers from one queue, and keeps what they submit from each set apart: labeled train images, labeled validation images and labeled test images each land in a dataset of their own.
Each labeler works through the sets in order: train first, then validation, then test. Set any one, two or all three, and the node returns a labeled dataset for each one you set. Labelers see a single Labeled button and never pick where an image goes. Use it to label the splits a model trains and is scored on, where mixing them would let a model be scored on images it trained on. For custom status buttons or a single source, use Labeling instead.
Once every required input is set, the node runs itself every 15 minutes to top up queues and take back unfinished work.
How it works
Checks each set you picked and the Classes ontology, and creates a labeled dataset for each set on its first run, reusing it on every later one.
Gives each person in Labelers a private queue, and sends the queue of anyone removed from the list back to the set it came from.
FAQ
Why did the run fail?
The run fails when none of Train images, Validation images or Test images is set, when a set no longer exists, or when Classes names no ontology, since every queue is pinned to one.
What happens when images are added to a set that was already done?
A queue that is empty goes back to the first set, in order, that has images again. A queue still holding images of a later set finishes those first.
When does unfinished work go back?
Images that sit in a queue longer than Return unclaimed work after (minutes) go back to the set they came from on the node’s next run, along with any annotation the labeler saved but never submitted.
What happens if I clear one of the sets while labelers work on it?
On the next run, a queue on that set gives back what it holds and moves to the first set still picked. The labeled dataset of the cleared set keeps what was already submitted.
What happens when I delete the node?
Every image still in a labeler’s queue goes back to the set it came from. The labeled datasets keep what was already submitted.
Can labelers add classes or use several status buttons?
No. Each queue offers one Labeled button, and the vocabulary is the one picked in Classes. Use Labeling for either.
Inputs
Images to label for training. Labelers get these first. Anything left unsubmitted comes back here. Key: TRAIN_DATASET.
Images to label as the validation set. Labelers get these once the train images run out. Key: VAL_DATASET.
Images to label as the test set. Labelers get these last. Key: TEST_DATASET.
Workspace members who label. Each gets one private queue that works through the splits in order: train, then validation, then test. Key: LABELERS.
Vocabulary every split is labeled in. Defaults to the train images’ ontology. Tick only part of it to hide the rest from labelers, or tick nothing to label in all of it. Values come from TRAIN_DATASET. Key: CLASSES.
Number of images a labeler holds at once. Their queue is topped up to this after each submit and on every node run. Minimum: 1. Key: BATCH_SIZE.
Images left unsubmitted in a labeler’s queue longer than this go back to the split they came from, so nobody sits on work they never finished. Minimum: 1. Key: CLAIM_TTL_MINUTES.
Outputs
Internal map of user IDs to their hidden labeling queue datasets. Hidden from the flow editor. Key: USERS_DATASETS.
Internal list of the datasets submitted images land in. Hidden from the flow editor. Key: OUTCOME_DATASETS.
Dataset the labeled train images land in. Empty when no train images are set. Key: LABELED_TRAIN.
Dataset the labeled validation images land in. Empty when no validation images are set. Key: LABELED_VAL.
Dataset the labeled test images land in. Empty when no test images are set. Key: LABELED_TEST.
Compact per-labeler queue status table object. Hidden from the flow editor. Key: STATUS.
JSON config
Machine-readable node interface for automation and advanced usage.