Why can't we report on the fields we have been filling in for two years?
Free text and unmanaged picklists. A field that accepts anything will collect fifty spellings of six things, and a picklist anyone can extend grows options that mean the same thing. Reporting then needs a manual cleanup pass every time, so eventually nobody reports on it. The fix is a controlled list with a named owner, a one-time mapping of historical values, and a separate note field so people stop inventing values in the main one.
Two drift mechanisms, two different fixes
Typing drift comes from free-text fields. Ten people describe the same thing ten ways, and none of them are wrong. There is no cleanup that prevents it from happening again, because the field itself is the problem.
Option sprawl comes from picklists that any user can extend. It looks controlled and is not. Over two years the list grows from eight options to sixty, several of which are near-synonyms added by someone who could not find the existing one.
The diagnostic takes five minutes
Run a distinct-value count with frequencies on every field you expect to report on. The shape of the output tells you which problem you have.
A long tail of values used once or twice is typing drift or option sprawl. A short list where one option holds an overwhelming share is a field people are defaulting through without thinking. And a large share sitting in Other means the list does not match reality, which is a design problem rather than a discipline problem. Each shape calls for a different response, and lumping them together as bad data guarantees the wrong fix.
Fix the list before you clean the data
Cleaning historical values while the field is still open is wasted work, because drift resumes the day after you finish. Close the list first, assign an owner who is the only person who can add options, and give people a note field for the cases the list does not cover.
Then read the note field monthly. It is the backlog of options you are missing, and it is far better evidence for changing the list than anyone's opinion in a meeting.
Where AI helps and where it does not
Mapping two years of messy historical free text onto a clean controlled vocabulary is genuinely good work for a language model. It handles spelling variation, abbreviations and phrasing differences that rules-based cleanup misses, and it can flag the values it was unsure about for a person to decide. That is a real use of operational AI on data you already own.
What it cannot do is keep the field clean going forward, because that is a process and permissions question rather than a modeling one. If the field stays open, you will be re-running the cleanup next year. Deciding which fields are controlled, who owns them, and how new options get added is the unglamorous half of making customer data reportable.
Topics: data hygiene · picklists · reporting · CRM fields
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.