Data-Centric AI & Annotation
Dataset quality as the object of study rather than a preliminary step. Evidenced by the Data Annotation Analyst role at Rooya AI, the CVAT annotation portfolio, and dataset analysis work at AMIRL.
Dataset quality treated as the object of study rather than the step before it. The reason is practical: across these projects the labels were the scarce resource, and the difference between a result and an artefact was usually a decision about the data rather than about the model.
The work has an applied half and a methodological one. Applied: annotating and validating industrial data as a data annotation analyst, and a CVAT portfolio covering bounding boxes, polygons, segmentation masks, keypoints and object tracking, with multi-class labelling and annotation quality control across it. Methodological: dataset analysis and evaluation design inside the research projects — patient- and case-level splitting, class balance, and stating when a test split is too small to support a claim.
The two halves meet at label consistency. A guideline that two annotators read differently sets a ceiling no architecture can lift, and it does not appear anywhere in the metrics.