01 · Propose
Let the model create the initial annotations
Generate an initial set of annotations, then review, adjust, and approve the results directly in the workspace—without switching between tools.
Browse the model zoo
Trusted by researchers and teams worldwide
MIT
Cambridge
Adelaide
NTU
Tsinghua
PKU
HIT
ZJU
CAS
IAEA
Tencent
DJI
Alibaba
AfricaMuseum
ColumbiaA complete annotation loop
Models produce the initial annotations, people review and refine them, and the resulting data moves directly into training or downstream workflows.
01 · Propose
Generate an initial set of annotations, then review, adjust, and approve the results directly in the workspace—without switching between tools.
Browse the model zoo
02 · Review
Review model-generated results, correct locations and classes, add missed objects, and make every annotation complete, consistent, and ready for training.
Read the user guide03 · Deliver
Turn reviewed annotations into training data, then bring improved models back into review to keep the data flywheel moving.
Read the training guideMultimodal annotation workspace
Unify tasks, annotation methods, model backends, and data formats across the complete workflow from data preparation to model application.
Multimodal tasks
Classification, detection, segmentation, pose, tracking, OCR, document parsing, video classification, captioning, VQA, multimodal conversations, and more.
Annotation geometry
Polygons, rectangles, cuboids, rotated boxes, circles, lines, points, masks, and task-specific shapes.
Model library
Integrate mainstream model families including YOLO, SAM, DINO, Qwen, and PPOCR into annotation workflows.
Inference stack
Use ONNX Runtime, TensorRT, OpenCV DNN, or PyTorch, with remote services such as SGLang, vLLM, and TGI.
Open formats
Work with COCO, VOC, YOLO, DOTA, MOT, masks, PPOCR, MM-Grounding, ShareGPT, and more.
Multilingual
Use X-AnyLabeling in English, Simplified Chinese, Japanese, or Korean.
Recently added
Purpose-built tools for document parsing, video classification, and multimodal conversations keep complex tasks moving in one workspace.
Document parsing
Parse layouts, tables, formulas, and text with PaddleOCR, then review every result in place.
See the document workflowVideo classifier
Mark frame-accurate segments, assign classes, review descriptions, and export clips or raw frame sequences.
See the video workflowChatbot
Work with vision-language models beside the current image and preserve approved responses as training data.
See the chatbot workflowAnnotate, review, and export locally, then connect the models, inference services, and open formats that fit your existing toolchain.
