Data labeling & annotation · CVAT.ai (open-source project, originated at Intel)
CVAT
Open-source, self-hosted tool focused on image and video annotation for computer vision, with bounding boxes, polygons and tracking.
CVAT (Computer Vision Annotation Tool) is an open-source annotation tool specialized for computer-vision tasks: bounding boxes, polygons, polylines, points, cuboids and skeletons on images, plus interpolation-based tracking across video frames so objects don't need to be re-drawn on every frame. It originated at Intel and is now maintained as an independent open-source project, deployable via Docker on a team's own infrastructure, with a hosted SaaS version (cvat.ai) also available for teams that don't want to self-host. AI-assisted labeling can be added through integrated model-serving (e.g., Segment Anything) to speed up annotation. Its scope is narrower than general multimodal platforms in this category — it is built specifically for vision data rather than text, audio or documents.
At a glance
| Vendor | CVAT.ai (open-source project, originated at Intel) |
|---|---|
| Pricing model | Open source + paid options |
| Free tier | Yes |
| Deployment | Self-hosted, Cloud |
| Open source | Yes (MIT) |
| Best for | Computer-vision teams that want a free, self-hosted tool focused specifically on image/video annotation. |
Pricing
The open-source core is free and self-hosted; a hosted SaaS/enterprise version with paid plans is also offered.
Pricing has not been verified yet — see the vendor's site.
Features
- Bounding box, polygon, polyline, point, cuboid and skeleton annotation
- Frame-interpolation tracking for video annotation
- AI-assisted labeling via integrated model serving (e.g., Segment Anything)
- Self-hosted deployment via Docker
- Hosted SaaS version (cvat.ai) available
- Multi-user projects with task assignment and review stages
- Import/export in common CV formats (COCO, YOLO, Pascal VOC)
Integrations
Profile last reviewed September 21, 2026