To become more productive across the coatings value chain
An AI agent does not know what a product name means. It cannot do anything with "Mergisol ME 107". But give it a glass transition temperature, an epoxy equivalent weight or a Hansen parameter, and it can start to reason and to predict. So if we want AI to help us formulate, we have to describe our raw materials as numbers and not as names. The trouble is that we don't have much of that data in coatings, and a lot of what we do know sits in the heads of a few experts. That is why I wanted to spend this workshop on a practical question. What data do we actually need across the value chain, and how could we make it available to each other.
Before working out which data we need, I asked everyone what they would actually do with it. People want to optimise and stay competitive, to run R&D faster and with less guesswork, to catch mistakes before they turn into complaints, and to build toward more sustainable products.
Optimise and stay competitive
Run R&D faster
Avoid mistakes and complaints
Predict, discover, do it sustainably
I opened with four epoxy diluents and one simple question. Which one reduces the viscosity of the system most at a 20% usage level? With only the four product names in front of you, this is almost impossible to answer unless you happen to have worked with all four yourself. And that is exactly the point. The name on its own tells an AI agent, or a new colleague, nothing it can work with.
Now look at the same four diluents with their properties next to them.
| Diluent | Chemistry | EEW | Funct. | Neat visc. mPa·s | Density g/cm³ | Blend at 20% mPa·s |
|---|---|---|---|---|---|---|
| Diluent E | glycidyl ether (cardanol) | 500 | 1 | 50 | 0.97 | 1500 |
| Diluent F | cardanol | n/a | 0 | 55 | 0.93 | 1750 |
| Diluent G | glycidyl ether (C12-C14) | 285 | 1 | 7.5 | 0.90 | 770 |
| Diluent H | epoxidized fatty acid ester (methyl) | 245 | 0 | 18 | 0.935 | 1229 |
With the chemistry, the epoxy equivalent weight, the functionality, the neat viscosity and the density in front of you, the question becomes answerable. Diluent G, the low-viscosity reactive glycidyl ether, brings the blend viscosity down the most, and you can read that off the numbers before running a single trial.
In the first breakout, each group walked down the value chain and answered two questions for every stage. What does this stage need from the others, and what could it provide in return? I deliberately left cost, IP and willingness to share out of it. For now the only question was what is technically possible and what already exists.
Needs from others
Can provide
Needs from others
Can provide
Needs from others
Can provide
One counterpoint from the audience about sharing models was to share the raw data instead. The case was that a plain record like "at this concentration of substance X, I measured this hardness," together with the metadata on how it was measured, lets the next company in the chain train its own model on numbers it can actually see, rather than trusting a black box it cannot look into.
In the second breakout, each group worked through how to make this data move between companies, and four approaches came up.
Through the TDS and MSDS they already send, an extended MSDS with more data points, or a certificate of analysis per batch on a shared platform.
What blocks it today
As part of incoming-goods control when the material arrives, or by pulling what already exists out of the ERP system.
What blocks it today
Pay an independent service to record and manage the data, then share either the raw data or the trained models through a shared platform.
What blocks it today
Bring the up and downstream steps you need inside the company and generate the data yourself.
What blocks it today
One idea we kept circling back to was that each company trains its own model and then shares the model, not the underlying data. Someone in the audience made a good counterpoint. Maybe we should share the raw data instead. A simple record like "at this concentration of substance X, I measured this hardness" lets everyone train their own model on data they can actually look at, rather than trusting a black box they cannot see into. I think they are right.
Whichever way you share, models or raw data, you hit the same requirement. Everyone has to describe exactly how the data was measured, ideally in one shared reference system. And that is the hard part. Right now everyone measures a little differently, and the moment the methods drift apart the data stops being comparable.
The idea that got the most heads nodding was an independent raw material database, run by a neutral party rather than a large supplier. Something like Evonik's Coatino already exists, but in the end Evonik wants to sell Evonik products, so there is a built-in bias. What people actually want is a neutral, maybe even open database that any formulator can use, holding raw materials as data rather than as PDF datasheets, ideally with a way to order samples directly from it. I think this could be genuinely valuable, and it is something I want to keep working on.
Through HOBUM and Experimental Design Hub I want to keep working on this. The first step I would like to take with a few partners is to work out exactly which data we need to improve which task. The second is to figure out how best to make that data available. Ideally that means a coating manufacturer together with someone on the application side. If that sounds interesting to you, get in touch.