> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.ovalbee.com/node-library/nodes/label-labeling-splits/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.ovalbee.com/_mcp/server. # Label train, validation and test > Hand a train, a validation and a test set to labelers from one queue, split by split, and keep each split's labeled images in a dataset of its own. ## Overview Hands the images of a train, a validation and a test set to labelers from one queue, and keeps what they submit from each set apart: labeled train images, labeled validation images and labeled test images each land in a dataset of their own. Each labeler works through the sets in order: train first, then validation, then test. Set any one, two or all three, and the node returns a labeled dataset for each one you set. Labelers see a single **Labeled** button and never pick where an image goes. Use it to label the splits a model trains and is scored on, where mixing them would let a model be scored on images it trained on. For custom status buttons or a single source, use [Labeling](/node-library/nodes/label-labeling) instead. Once every required input is set, the node runs itself every 15 minutes to top up queues and take back unfinished work. ```mermaid flowchart LR split["Train/validation split"] --> label["Label train, validation and test"] --> train["Train"] ``` ## How it works Checks each set you picked and the **Classes** ontology, and creates a labeled dataset for each set on its first run, reusing it on every later one. Gives each person in **Labelers** a private queue, and sends the queue of anyone removed from the list back to the set it came from. Tops each queue up to **Batch size** from the set it is on. A queue starts on the first set you picked and stays on a set until that set has no images left and the labeler has submitted everything they hold. Moves an image and its annotation to its set's labeled dataset as soon as the labeler submits it, then refills the queue right away, from the next set once this one runs out. ## FAQ #### Why did the run fail? The run fails when none of **Train images**, **Validation images** or **Test images** is set, when a set no longer exists, or when **Classes** names no ontology, since every queue is pinned to one. #### What happens when images are added to a set that was already done? A queue that is empty goes back to the first set, in order, that has images again. A queue still holding images of a later set finishes those first. #### When does unfinished work go back? Images that sit in a queue longer than **Return unclaimed work after (minutes)** go back to the set they came from on the node's next run, along with any annotation the labeler saved but never submitted. #### What happens if I clear one of the sets while labelers work on it? On the next run, a queue on that set gives back what it holds and moves to the first set still picked. The labeled dataset of the cleared set keeps what was already submitted. #### What happens when I delete the node? Every image still in a labeler's queue goes back to the set it came from. The labeled datasets keep what was already submitted. #### Can labelers add classes or use several status buttons? No. Each queue offers one **Labeled** button, and the vocabulary is the one picked in **Classes**. Use [Labeling](/node-library/nodes/label-labeling) for either. ## Inputs **`Train images`** `dataset` Images to label for training. Labelers get these first. Anything left unsubmitted comes back here. Key: `TRAIN_DATASET`. --- **`Validation images`** `dataset` Images to label as the validation set. Labelers get these once the train images run out. Key: `VAL_DATASET`. --- **`Test images`** `dataset` Images to label as the test set. Labelers get these last. Key: `TEST_DATASET`. --- **`Labelers`** `users` — required Workspace members who label. Each gets one private queue that works through the splits in order: train, then validation, then test. Key: `LABELERS`. --- **`Classes`** `ontology` — required Vocabulary every split is labeled in. Defaults to the train images' ontology. Tick only part of it to hide the rest from labelers, or tick nothing to label in all of it. Values come from `TRAIN_DATASET`. Key: `CLASSES`. --- **`Batch size`** `integer` — required, default: 10 Number of images a labeler holds at once. Their queue is topped up to this after each submit and on every node run. Minimum: `1`. Key: `BATCH_SIZE`. --- **`Return unclaimed work after (minutes)`** `integer` — required, default: 120 Images left unsubmitted in a labeler's queue longer than this go back to the split they came from, so nobody sits on work they never finished. Minimum: `1`. Key: `CLAIM_TTL_MINUTES`. --- ## Outputs **`User queues`** `object` Internal map of user IDs to their hidden labeling queue datasets. Hidden from the flow editor. Key: `USERS_DATASETS`. --- **`Outcome datasets`** `object` Internal list of the datasets submitted images land in. Hidden from the flow editor. Key: `OUTCOME_DATASETS`. --- **`Labeled train`** `dataset` Dataset the labeled train images land in. Empty when no train images are set. Key: `LABELED_TRAIN`. --- **`Labeled validation`** `dataset` Dataset the labeled validation images land in. Empty when no validation images are set. Key: `LABELED_VAL`. --- **`Labeled test`** `dataset` Dataset the labeled test images land in. Empty when no test images are set. Key: `LABELED_TEST`. --- **`Status`** `object` Compact per-labeler queue status table object. Hidden from the flow editor. Key: `STATUS`. --- ## JSON config Machine-readable node interface for automation and advanced usage. #### Show JSON ```json { "name": "Label train, validation and test", "description": "Hand a train, a validation and a test set to labelers from one queue, split by split, and keep each split's labeled images in a dataset of its own.", "category": "Label", "namespace": "ovalbee", "templateKey": "label/labeling_splits", "version": "v1", "inputs": [ { "key": "TRAIN_DATASET", "label": "Train images", "type": "dataset", "description": "Images to label for training. Labelers get these first. Anything left unsubmitted comes back here.", "required": false, "default": null, "visibleWhen": null, "options": { "creatable": false } }, { "key": "VAL_DATASET", "label": "Validation images", "type": "dataset", "description": "Images to label as the validation set. Labelers get these once the train images run out.", "required": false, "default": null, "visibleWhen": null, "options": { "creatable": false } }, { "key": "TEST_DATASET", "label": "Test images", "type": "dataset", "description": "Images to label as the test set. Labelers get these last.", "required": false, "default": null, "visibleWhen": null, "options": { "creatable": false } }, { "key": "LABELERS", "label": "Labelers", "type": "users", "description": "Workspace members who label. Each gets one private queue that works through the splits in order: train, then validation, then test.", "required": true, "default": null, "visibleWhen": null }, { "key": "CLASSES", "label": "Classes", "type": "ontology", "description": "Vocabulary every split is labeled in. Defaults to the train images' ontology. Tick only part of it to hide the rest from labelers, or tick nothing to label in all of it.", "required": true, "default": null, "visibleWhen": null, "options": { "ref": "TRAIN_DATASET", "selectable": true, "creatable": true } }, { "key": "BATCH_SIZE", "label": "Batch size", "type": "number", "description": "Number of images a labeler holds at once. Their queue is topped up to this after each submit and on every node run.", "required": true, "default": 10, "visibleWhen": null, "options": { "type": "integer", "min": 1 } }, { "key": "CLAIM_TTL_MINUTES", "label": "Return unclaimed work after (minutes)", "type": "number", "description": "Images left unsubmitted in a labeler's queue longer than this go back to the split they came from, so nobody sits on work they never finished.", "required": true, "default": 120, "visibleWhen": null, "options": { "type": "integer", "min": 1 } } ], "outputs": [ { "key": "USERS_DATASETS", "label": "User queues", "type": "object", "description": "Internal map of user IDs to their hidden labeling queue datasets.", "artifact": false, "badge": false, "badgeOpens": null, "hidden": true }, { "key": "OUTCOME_DATASETS", "label": "Outcome datasets", "type": "object", "description": "Internal list of the datasets submitted images land in.", "artifact": false, "badge": false, "badgeOpens": null, "hidden": true }, { "key": "LABELED_TRAIN", "label": "Labeled train", "type": "dataset", "description": "Dataset the labeled train images land in. Empty when no train images are set.", "artifact": false, "badge": false, "badgeOpens": null, "hidden": false }, { "key": "LABELED_VAL", "label": "Labeled validation", "type": "dataset", "description": "Dataset the labeled validation images land in. Empty when no validation images are set.", "artifact": false, "badge": false, "badgeOpens": null, "hidden": false }, { "key": "LABELED_TEST", "label": "Labeled test", "type": "dataset", "description": "Dataset the labeled test images land in. Empty when no test images are set.", "artifact": false, "badge": false, "badgeOpens": null, "hidden": false }, { "key": "STATUS", "label": "Status", "type": "object", "description": "Compact per-labeler queue status table object.", "artifact": false, "badge": false, "badgeOpens": null, "hidden": true } ], "automation": { "requires_configuration": true, "schedule": "*/15 * * * *", "discard_history": "on_success" }, "widgets": { "widget": { "id": "table", "settings": { "data": { "type": "variable", "value": "self.outputs.STATUS" } } }, "secondaryWidget": { "id": "open-labeling-tool", "settings": { "name": { "type": "input", "value": "Open annotation tool" } } } } } ``` > Hand a train, a validation and a test set to labelers from one queue, split by split, and keep each split's labeled images in a dataset of its own.