Josef Hainz
“High risk” means: retiring within 24 months and at least 8 h of downtime on faults the shift could only call this person for.
From fault tickets every company already has, code counts in milliseconds who knowledge depends on and what that costs. Only that is what the foreman gets asked, about patterns rather than dates. His answer then shows up in the ticket, with its source.
Brandl Antriebstechnik GmbH is made up, the same company as in project 01. People and documents are made up. Not a maintenance manual.
Every company has fault tickets. Who was called, whether a cause is written down, how long the machine stood and at what time: code counts that. No interview, no AI model.
That's how it starts with a client: “Send me an export of your fault tickets. Tomorrow you'll see who your plant depends on.”
“High risk” means: retiring within 24 months and at least 8 h of downtime on faults the shift could only call this person for.
Selin, Markus, Robert, Dana, Tobias: no fault that only they could be called for.
5 red points, sorted by downtime. The question starts with the fault as the shift knows it, then how often and what has been tried. Nobody remembers one night three years ago, but everyone remembers a pattern.
If a ticket only says “running again”, exactly one short question goes out the next morning to whoever fixed it. A template in code, no model.
Yesterday 15:30, K-2: “Overtemp. shutdown”. Selin called you, then it ran again. What was it?
The same fault ticket, the same search, the same model. The plain AI chatbot reads all documents as text chunks, “with follow-up” reads the same documents ordered by machine and fault, plus the approved knowledge from the follow-up. Measured on 30 test cases, scored by code.
“HA1 cooling water flow min, hardening down. Night shift started up after weekend”
Correct would be: HA-1 is on cooling circuit KK-2, and that has air in it after the weekend because P-204 draws air: bleed P-204 first, then reset HA-1.
Check strainer and cooling water supply
On the alarm 'Cooling water flow min', the manual says to check the strainer and the cooling water supply. In the past this fault was often cleared just by acknowledging it, with downtimes between 20 and 35 minutes. Also check whether the flow switch is defective.
Bleed P-204, then acknowledge HA-1
Since the fault occurs on start-up after the weekend, bleed pump P-204 before acknowledging HA-1. Note that the manual generally calls for checking the strainers and the cooling water supply, while the specific instruction for after the weekend puts bleeding P-204 first.
Check strainer KK-2 and the cooling water supply. Acknowledge the fault at the operator panel only after it has been fixed.
On the alarm 'Cooling water flow min', check strainer KK-2 and the cooling water supply. Once the cause is fixed, acknowledge the fault at the operator panel. Note that the records contradict each other on this fault, as simply replacing the flow switch did not fix it permanently in the past.
30 cases, each sent to three systems. The chatbot sees all documents as text chunks, “documents only” sees the same documents ordered by machine and fault, “with follow-up” also sees Josef's approved answers.
Recent fault tickets, kept out of the documents before the run. 5 of them can only be solved with follow-up: A 0 · B 0 · C 3.
Fault tickets written by a second agent that knew neither the rules, the prompts nor the test set. Reported separately.
A number without a source never reaches the ticket. Code checks it, not a model.
In the sample plant, reported at night 6 times and fixed only after 6 am: 37 hours of downtime.
Costs from the recorded run. System with follow-up, $1 = €0.86. Not included: the link to the ticket system and the time for approvals, which is the real effort. You set the hourly rate; it is not my estimate.
It doesn't ask at random. The knowledge check shows where knowledge is missing and has already cost downtime. Only that gets asked, sorted by downtime.
The questions are about patterns, not dates. What he no longer knows stays open, and the system then says “not written down anywhere” instead of guessing. Anything new gets asked the next morning.
Josef Hainz is played by Gemini 3.1 Flash-Lite. It knows his fact sheet of 14 facts completely, and I wrote both the sheet and the test cases. A real foreman remembers with more gaps. This has not been tested with a real maintenance technician. I don't play Josef myself either, because I'm not a maintenance technician.
It only finds knowledge that left traces in the documents: who was called, where it only says “running again”. What was never reported, or whoever nobody writes down, it doesn't see.
An evaluation per person is capable of recording performance and behaviour, and therefore typically subject to co-determination in Germany. In practice only with a works agreement, for knowledge capture, never for performance reviews.
Follow-up takes the foreman's time. Planning it is the lead's job. The tool only makes it as short as possible: here 5 topics instead of a day of interviews.
Company, machines, people and all 200 documents are made up. Real fault tickets are messier, and the rules would need tuning for each plant.
30 test cases and 8 outside test cases, fixed before the run, scored by code. That is strict and can count a correct answer in other words as wrong. After the first run I fixed two bugs in the scorer, the same for all systems. Every number comes from one run: a difference below two cases is a tendency.
Recorded on 04/10/2026. Only code runs live in your browser: the knowledge check and the rule for fresh questions. All model output is recorded; a visit costs nothing. Documents up to 07/2026. Models: Reading Gemini 3.1 Flash-Lite · Verdicts Gemini 3.8 Flash · Follow-up Gemini 3.8 Flash · J. Hainz Gemini 3.1 Flash-Lite · Answers Gemini 3.1 Flash-Lite · Embeddings Gemini Embedding 2.