Start with process samples, not a perfect lake. Two to four weeks of tickets, call notes, order files, or warehouse events give us the real distribution of mess. Include happy paths and the ugly corner cases staff swear are rare. Metadata on volumes per day and current error rates set the scoreboard.
System catalogs come next. Names, owners, auth methods, rate limits, and existing API docs prevent surprises. If an API is missing, screen flows or SFTP drops still work with more risk. Sandbox credentials long before coding starts save everyone. Virginia teams with lean IT should assign a single coordinator for access requests.
Policy documents ground conversational systems. FAQs, SOPs, refund rules, and care scripts feed retrieval stores so answers stay approved. Without them, models improvise and compliance loses. Versioned PDFs with clear owners beat tribal knowledge on sticky notes.
Privacy notes matter early. Flag fields that are health, financial, or student related. We design redaction and retention around those tags rather than bolt them later. If prior anonymization work is available, share it. Learning from law-enforcement redaction pipelines shows how strict paths still leave analytics useful.
You do not need labeled machine learning gold on day one. Rules plus light models often win first. Labels grow from override logs during assist mode. That approach respects how mid-market Harrisonburg data actually arrives.