We need sample exports of leads, contacts, accounts, opportunities, and activities with field definitions. Two to four weeks of recent history is enough for a pilot model, provided it covers both wins and losses. Consent and do-not-contact flags must travel with the export. Identity rules and ownership maps matter more than sheer volume. Systems diagrams listing mail, calendar, support, and billing connections complete the starter kit.
Quality beats quantity. Ten thousand clean rows outperform a million rows with broken emails and blank companies. We also need access policies: who may approve production writes, which fields are forbidden from model context, and where secrets live. For healthcare-adjacent or student-adjacent data around UVA, purpose limitation notes should arrive in week one. Startup ecosystems in Charlottesville and the broader central Virginia tech community often lack perfect warehouses. That is fine. Structured CRM tables plus labeled outcomes are enough if definitions are shared.
If data is messy, Discovery expands into a cleanup phase rather than pretending. We never invent synthetic customers to flip a demo. You should expect a data readiness memo before any paid model training. When exports are blocked by legal review, we work from scrubbed samples under an agreed protocol so project time doesn't stall. Clear data inputs are the single best predictor of on-time AI CRM Automation delivery.