ExperimentalDesignHub Workshop recap

What data do we need?

To become more productive across the coatings value chain

SPFA Workference 2026 · Workshop 2 · 18 June 2026 · Online

An AI agent does not know what a product name means. It cannot do anything with "Mergisol ME 107". But give it a glass transition temperature, an epoxy equivalent weight or a Hansen parameter, and it can start to reason and to predict. So if we want AI to help us formulate, we have to describe our raw materials as numbers and not as names. The trouble is that we don't have much of that data in coatings, and a lot of what we do know sits in the heads of a few experts. That is why I wanted to spend this workshop on a practical question. What data do we actually need across the value chain, and how could we make it available to each other.

What we would do with the data

Before working out which data we need, I asked everyone what they would actually do with it. People want to optimise and stay competitive, to run R&D faster and with less guesswork, to catch mistakes before they turn into complaints, and to build toward more sustainable products.

Optimise and stay competitive

optimizationcompetitive adjust formulationsproduct optimization manufacturing optimization

Run R&D faster

faster R&Dfaster development more efficient searchdata reuse learn faster

Avoid mistakes and complaints

quality controlavoid mistakes claim prevention

Predict, discover, do it sustainably

predict outcomesdiscover unknowns understand the productsustainable products sustainable R&D

The thought experiment we started with

I opened with four epoxy diluents and one simple question. Which one reduces the viscosity of the system most at a 20% usage level? With only the four product names in front of you, this is almost impossible to answer unless you happen to have worked with all four yourself. And that is exactly the point. The name on its own tells an AI agent, or a new colleague, nothing it can work with.

Now look at the same four diluents with their properties next to them.

DiluentChemistryEEWFunct. Neat visc.
mPa·s
Density
g/cm³
Blend at 20%
mPa·s
Diluent Eglycidyl ether (cardanol)5001500.971500
Diluent Fcardanoln/a0550.931750
Diluent Gglycidyl ether (C12-C14)28517.50.90770
Diluent Hepoxidized fatty acid ester (methyl)2450180.9351229

With the chemistry, the epoxy equivalent weight, the functionality, the neat viscosity and the density in front of you, the question becomes answerable. Diluent G, the low-viscosity reactive glycidyl ether, brings the blend viscosity down the most, and you can read that off the numbers before running a single trial.

Breakout 1 — What data do we need?

In the first breakout, each group walked down the value chain and answered two questions for every stage. What does this stage need from the others, and what could it provide in return? I deliberately left cost, IP and willingness to share out of it. For now the only question was what is technically possible and what already exists.

Raw material supplier

Needs from others

  • What the product is used for and the resistance it has to meet
  • Regulatory restrictions, up and down the chain
  • Honest feedback from the user
  • Stable quality and availability from its own suppliers
  • A model that maps a raw material's features to its performance

Can provide

  • Composition and raw data (chemical type, ingredients, concentrations, solid content, density)
  • Physicochemical profiling (Hansen parameters, particle size, molecular weight distribution, surface chemistry)
  • How each property was measured
  • A trained model on its own raw materials that could speed up experimentation

Formulator

Needs from others

  • Raw material data from the supplier, with its provenance
  • A model linking a raw material's features to performance, on a standardised basis
  • Application conditions from the applicator (line, drying, storage, part geometry)
  • Clear property targets (viscosity, glass transition, MFT, tensile strength)
  • Feedback from the user

Can provide

  • Test results, the failures as well as the passes
  • Stability and compatibility data
  • Processing instructions and ease of production
  • Application parameters for the applicator

Applicator

Needs from others

  • Any change in the formulation
  • Application and curing parameters
  • The standards and end use it has to meet

Can provide

  • Usability feedback from the floor
  • Conformance to standards and certifications
  • Input and output models of its own process

One counterpoint from the audience about sharing models was to share the raw data instead. The case was that a plain record like "at this concentration of substance X, I measured this hardness," together with the metadata on how it was measured, lets the next company in the chain train its own model on numbers it can actually see, rather than trusting a black box it cannot look into.

What I took from this round. For coating raw materials, what would really help is a shared, open source database. Raw materials held as data, run by a neutral party rather than a big supplier with its own products to sell. That is the thing I think would be worth building.

Breakout 2 — How do we make the data available?

In the second breakout, each group worked through how to make this data move between companies, and four approaches came up.

1The supplier shares it

Through the TDS and MSDS they already send, an extended MSDS with more data points, or a certificate of analysis per batch on a shared platform.

What blocks it today

2We measure it ourselves

As part of incoming-goods control when the material arrives, or by pulling what already exists out of the ERP system.

What blocks it today

3A neutral third party measures it

Pay an independent service to record and manage the data, then share either the raw data or the trained models through a shared platform.

What blocks it today

4We build it in-house

Bring the up and downstream steps you need inside the company and generate the data yourself.

What blocks it today

The question is still too broad to act on. In the first breakout we brainstormed what data could be shared, and we touched on why we would want to share it. But none of it is specific enough yet. The next step is to narrow it down. Which data exactly, shared with whom, and what each side gets out of it. Once that is clear, the question of how to make the data available becomes much easier to answer.

What I took away

Share the models, or share the raw data

One idea we kept circling back to was that each company trains its own model and then shares the model, not the underlying data. Someone in the audience made a good counterpoint. Maybe we should share the raw data instead. A simple record like "at this concentration of substance X, I measured this hardness" lets everyone train their own model on data they can actually look at, rather than trusting a black box they cannot see into. I think they are right.

Standardisation is the recurring requirement

Whichever way you share, models or raw data, you hit the same requirement. Everyone has to describe exactly how the data was measured, ideally in one shared reference system. And that is the hard part. Right now everyone measures a little differently, and the moment the methods drift apart the data stops being comparable.

A neutral raw material database

The idea that got the most heads nodding was an independent raw material database, run by a neutral party rather than a large supplier. Something like Evonik's Coatino already exists, but in the end Evonik wants to sell Evonik products, so there is a built-in bias. What people actually want is a neutral, maybe even open database that any formulator can use, holding raw materials as data rather than as PDF datasheets, ideally with a way to order samples directly from it. I think this could be genuinely valuable, and it is something I want to keep working on.

Where we go from here

Through HOBUM and Experimental Design Hub I want to keep working on this. The first step I would like to take with a few partners is to work out exactly which data we need to improve which task. The second is to figure out how best to make that data available. Ideally that means a coating manufacturer together with someone on the application side. If that sounds interesting to you, get in touch.