Building a reliable segmentation dataset from thousands of non-standardized medical product images.
Overview
A European medical device manufacturer approached Plavno after its internal computer vision team encountered a data bottleneck.
The company had a large image collection of rigid and semi-rigid medical products — including devices such as catheters, infusion components, and moulded plastic parts — but the images had been captured in different environments and were not suitable for direct model training. Object boundaries were often affected by reflections off plastic and glass surfaces, shadows, cropping, low contrast, and partial visibility. The annotation itself was not especially complex on a single image, but repeating it consistently across thousands of files required a much stricter process.
Plavno worked with the client to formalize the segmentation rules, prepare a representative pilot set, and establish a production workflow with separate annotation and review stages. The dataset was processed in batches so that recurring errors could be identified early rather than discovered after the full delivery.
10,000 + medical product images processed through the annotation pipeline
Pixel-level masks created across varied image conditions, including reflective plastic/glass surfaces, low contrast, and partial occlusion
Automated technical validation applied to every submitted mask
Independent quality review separated from first-pass annotation
Ambiguous images routed through a documented adjudication process
Validated batches delivered in the client's required training-data format
Image-level review, correction, and approval history retained throughout the project
Client involvement focused on rules and exceptional cases rather than full-dataset manual review

The client had already accumulated enough source material to support the next stage of its computer vision product, but the image library had not been created under a single data collection standard.
The target product appeared under different lighting conditions, against different surfaces, and at different angles. Some images had clear boundaries, while others contained reflections off plastic or glass device surfaces, weak contrast, partial occlusions, cropping, or nearby objects with similar shapes and colors.
Before the data could be used for model training, the product had to be separated from the surrounding image through an accurate segmentation mask. The main difficulty was not drawing one acceptable outline. It was ensuring that the same interpretation was applied thousands of times.
Without detailed rules, two annotators could produce visually reasonable but inconsistent masks. One might include a narrow product component that another treated as background. Shadows and reflective edges — especially on clear or glossy plastic medical device surfaces — could also be handled differently depending on individual judgment.
The client's internal computer vision team could review representative examples and define the intended output, but it could not manually inspect the full dataset. The project therefore required a delivery partner capable of managing both annotation volume and quality assurance.

The initial instruction—separate the visible product from the background—was too broad for production. The team first had to define how to handle shadows, reflections, unclear edges, cropped products, detached components, thin structures, occlusions, and low-quality images.
The second challenge was maintaining consistent quality throughout a long annotation cycle. Continuous QA was required to identify technical errors, inconsistent boundaries, recurring mistakes, and ambiguous cases before they affected the wider dataset.

Solution
Plavno created a controlled annotation workflow with separate production, validation, QA, and escalation stages: Dataset assessment and guideline development, Pilot annotation and team calibration, Batch-based segmentation, Independent QA and validated delivery. A representative pilot dataset was used to establish project-specific rules and approved reference masks. During production, every mask passed automated validation and independent visual review. Ambiguous images were escalated for senior review instead of being resolved through individual assumptions.
Pixel-level segmentation of medical product images
Project-specific visual annotation guidelines
Representative pilot dataset before scale-up
Approved reference masks for team calibration
Separate annotation and QA responsibilities
Automated validation of every submitted mask
Independent visual review
Documented edge-case adjudication
Image-level status and correction tracking
Controlled batch delivery
Export in the client's required model-training format
Full review and approval history
Dataset Assessment: Images were grouped by boundary clarity, visibility, background complexity, reflections, cropping, resolution, and annotation difficulty. This helped estimate workload and identify images requiring additional review.
Guideline Development: The team documented correct and incorrect masks, inclusion rules, boundary treatment, reflections, shadows, cropped objects, and escalation criteria. The guideline was updated whenever new image patterns appeared.
Pilot and Calibration: A representative pilot batch was completed before production. It helped clarify missing rules, test annotator consistency, validate the workflow, and create approved reference examples. Annotators completed calibration before receiving production access.
Production Segmentation: Images were processed in controlled batches and assigned statuses such as ready for review, correction required, escalated, rejected, or approved.
Automated Validation: Every mask was checked for technical issues, including empty masks, invalid contours, missing annotations, disconnected fragments, incorrect object counts, and export errors.
Independent QA: Reviewers checked contour accuracy, missing product areas, background inclusion, reflections, shadows, cropping, and compliance with the latest guideline. Failed masks were returned for correction.
Edge-Case Adjudication: Ambiguous images were escalated to senior reviewers. Final decisions were documented, added to the guideline, and applied to similar images across the dataset.
Batch Delivery: Approved data was delivered in batches with segmentation outputs, QA statuses, revision records, exclusions, and validation results. This allowed the client to use completed data earlier and reduced the risk of late-stage systematic errors.
Consistent annotation process applied across a large-scale image dataset
Client's computer vision team not used as final QA department
Parallel annotation work by multiple team members
Centralized interpretation and resolution of edge cases
Continuous QA throughout the project, not just at the end
Early detection of recurring annotation issues
Controlled and documented rework process
Image-level traceability for every sample
Staged delivery of fully validated data batches
Client involvement in defining segmentation standards and product-specific rules
Plavno responsible for routine production and end-to-end quality control
Architecture Overview
Data Intake Layer: Images were normalized, assigned stable identifiers, and linked to metadata for tracking batch ownership, annotation progress, QA status, corrections, escalations, and delivery versions.
Annotation Layer: Annotators created pixel-level masks using approved guidelines and reference examples. The workspace supported contour refinement, brush and polygon tools, mask controls, task statuses, reviewer comments, and correction history.
Validation Layer: Automated checks verified mask integrity and export consistency before files entered QA or the client's model-training pipeline.
QA Layer: Independent reviewers compared each mask with the source image and current standards. Findings helped identify recurring errors, unclear rules, problematic image groups, and batches requiring additional review.
Governance Layer: Role-based workflows controlled annotation, review, correction, approval, and export. Every change—from the original mask to final approval—remained traceable.

Value
Maintaining consistent boundary interpretation across a large and visually diverse image dataset
Masks followed documented rules for product edges, narrow components, cropping, shadows, and reflections.
The person approving a mask was not the same person who created it.
Unclear images were escalated and documented instead of being resolved through individual assumptions.
Annotation changes and approval decisions remained linked to the original image.
Benchmarks
A production workflow designed for high-volume annotation without sacrificing review discipline
10,000+ Images Processed — the workflow supported a large dataset containing a broad range of image conditions
Technical checks were applied before human review.
The dataset was reviewed and delivered incrementally, making recurring issues easier to detect and contain.
Uncertain images followed one documented decision path instead of being interpreted independently.
Target inter-annotator boundary agreement (IoU) before release to production: commonly 0.85–0.90+ for well-defined object segmentation tasks Typical first-pass QA acceptance rate for a mature, calibrated annotation team on a well-specified task: roughly 85–95%, remainder returned for correction rather than rejected outright Typical throughput for pixel-level polygon/brush segmentation on medical product imagery: a few hundred to low thousands of images per annotator per week, depending on object complexity — team sized to the client's timeline
Data Protection
The exact client security configuration and infrastructure details remain confidential under the project NDA.
The project included:
Role-based access for annotators, reviewers, and project managers
Controlled project workspaces
Encrypted data transfer and storage
Image-level audit history
Restricted export permissions
Project-specific retention rules
Confidentiality agreements for team members
Documented dataset delivery and deletion procedures

Innovative Experience
Pixel-level annotation and QA workflows for medical products and other high-accuracy computer vision datasets

Eugene Katovich
Sales Manager
Plavno manages the complete segmentation workflow — from pilot preparation and annotation guidelines to independent review, automated validation, adjudication, and model-ready dataset delivery.
Discuss Your DatasetManaged Annotation Delivery
From a non-standardized image library to a reviewed segmentation dataset ready for computer vision training
Review the Dataset
Analyze image variation, identify difficult categories, and select a representative pilot set.
Define the Standard
Create visual segmentation rules, approve reference masks, and calibrate the annotation team.
Annotate and Validate
Produce masks in controlled batches and apply automated technical checks to every submission.
Review and Deliver
Run independent QA, adjudicate uncertain images, and deliver approved datasets with full traceability.
A consistent and traceable image dataset delivered without shifting the full QA burden to the client
The source image collection was converted into structured, reviewed training data.
The client's team focused on defining rules and resolving unusual cases rather than manually reviewing every mask.
Pilot calibration, visual guidelines, and independent QA reduced variation across annotators and batches.
Batch-based delivery allowed systematic issues to be corrected before they affected the entire dataset.
Every approved image retained a history of annotation, review, correction, and final acceptance.
Tools We Used
Core tools behind image annotation, automated validation, QA, and controlled dataset delivery
Project Estimator
The estimated time to launch the product
Clear vision of functionality you need
15% discount on your first sprint

Frequently Asked Questions
Common questions about medical image segmentation and annotation delivery
The project included more than 10,000 medical product images.
The team created pixel-level segmentation masks around the visible target product.
No. The client approved the annotation standard and supported exceptional decisions, while Plavno managed routine annotation, QA, corrections, and batch acceptance.
Every mask passed technical validation and an independent visual review. Unclear images were escalated through a documented adjudication workflow. See “Quality Benchmarks” above for the QA thresholds the process is designed around.
No. The dataset was processed and delivered in controlled batches so recurring issues could be found and corrected earlier.
Yes. Annotation outputs can be prepared in COCO JSON, polygon, PNG mask, or a client-specific schema, depending on the downstream training pipeline.
Related Projects
About Plavno

Senior engineers + proven AI components to accelerate time-to-value.

From MVPs to enterprise platforms at global scale.

From extension UX to GPU pipelines and global scale.
Testimonials
Contact Us
Plavno experts contact you within 24h
Discuss your project details
We can sign NDA for complete secrecy
Submit a comprehensive project proposal with estimates, timelines, team composition, etc
Plavno has a team of experts ready to start your project. Ask us!

Vitaly Kovalev
Sales Manager

Fill in the form
to access the case study.