When it comes to AI in the data classification universe, there will be no takeover of the machines. Well, not for a very long time. Don’t get me wrong, I agree that AI is an essential and valuable tool when it comes to spend data classification; we can now classify millions of data in a fraction of the time it would take to do manually. But this is only valuable if that data is classified correctly.
And how do we check that? We use humans.
Back to the beginning
Over the last few years, many of the AI systems used for spend data classification have been based around machine learning, and now, the hot topic of the moment, Gen AI. However, for success, the training data must be 100% accurate and completely clean. The only way to achieve this is through human involvement. I have worked for many years in spend data classification, and have even worked alongside AI to improve data. However, I still believe that to achieve the best results, you need people to initially classify the spend data. And that’s what we do for our clients.
People often assume that many spend analytics companies are using machines to classify their data. Yes, they can do a lot of the heavy lifting, but if you ask them what percentage of their data is being cleaned by machine learning or AI, I bet the percentage would be a lot lower than you think. If AI did everything, I wouldn’t have been running a successful manual spend data classification business for seven years!
If we had “nailed” AI spend data classification, there would be an industry leader selling their product to the market. However, as it currently stands, there are a number of companies using many different methods, scripts, and rules; all trying to perfect their product.
AI in spend data classification – why is it so difficult?
Why has nobody perfected this process yet? Well, unlike other industries where AI can completely replace human intervention, spend data classification is not a black and white, right or wrong process. There are context, variables and conditions, and these all influence the spend data classification.
I like to use DHL as an example. A professional services company is most likely to use them for courier services. In contrast, a manufacturer is more likely to use their supply chain, logistics and warehousing services. The AI would have to see the supplier and the industry the data comes from several times to learn and perfect this, helped of course by humans.
Here’s another example. Depending on how the machine learning works, it can misread invoice line descriptions. Something like “Taxi from hotel” can be confusing and the data could end up being classified as a hotel rather than a taxi. Of course, this would need to be rectified manually.
Another example could be a generic description of “cleaning”. It might seem like an obvious choice; however, depending on who the supplier is, it will depend on the service. We’ve seen this appear in IBM for keyboard or data cleaning, but the first assumption for that description would be janitorial services.
Let’s remember that one supplier can be many things, depending on the spend amount. If a hotel has a line value of up to around £5,000 then this is most likely to be for accommodation. A higher value could mean it’s for venue hire. However, it’s very much dependent on if there is an invoice line description that can act as a guide. This makes it very tricky to write rules for, or for the AI to learn from. That’s why we still need the human touch.
AI in spend data classification – a potential opportunity
Where AI can really shine through is for ever-changing items such as lab chemicals or electronic parts. What would take considerable time for a human to research, find and classify each time, can be classified by AI at lightning speed. This is of course after the initial manual classification…
What about Gen AI?
Many people are concerned about Gen AI taking their jobs and replacing them, however based on my experience, I don’t think this will be the case. We can use Gen AI to enhance our work life, and make it more efficient. Yet, we still need people to check the information we receive is correct and accurate, which as we know isn’t always the case. Gen AI can hallucinate; I’ve seen this for myself when verifying addresses of some schools.
When I put in the name of the school and the postcode, only 3 of the 15 entries were correct. Chat GPT had hallucinated a completely different first line of the address. If we appreciate that Gen AI works better in some areas than others, we can leverage this. My advice? Use it wisely, not blindly.
But this is not the first time we have been fearful of new technology. Many were concerned about RPA (Robotic Process Automation) when it launched. Organisations deployed this technology to improve extremely manual tasks such as document scanning, order processing, CRM updates and payroll processing.
While employees may initially have felt threatened by this new technology, the reality is that RPA has taken on many of the tasks we humans detest. They are manual, time-consuming and mundane tasks that can be processed efficiently, and more quickly and accurately than by humans. This is unlike AI for data classification which requires a heavy human involvement.
If introduced into the organisation correctly, employees will understand and appreciate that just like RPA, Gen AI is there to help free up their workload from these tasks. This allows them to focus on the areas that need more skill and are more rewarding.
What does the future hold for AI and spend data classification?
As data is ever-changing and updating, I can’t see a time when there won’t be a need for human involvement in spend data classification, even if this is simply to update new and unseen data. While full dependency on AI may work in other industries, I don’t think this approach will work when applied to spend data classification. The same applies to Gen AI; there will always be a need for human involvement at some stage.
Yet – I am no oracle and have not seen and experienced every AI tool available. I would happily be proved wrong. However, at this stage and for the coming years I am not concerned about my fate and the need for The Classification Guru’s services.
Enjoyed this post? Read this one next – Spoiler! AI & ChatGPT can co-exist with spend data classification

