Building in the age of AI · Evening #3
Getting an AI feature to work in a demo takes an afternoon now. Keeping it working in production is where the real effort goes. Outputs shift when a provider updates a model. Costs that looked trivial in a prototype grow with every new user. A prompt change fixes one case and quietly breaks three others, and nobody notices until a customer does.
Most teams find these problems after go-live, when the architecture is already set and the budget already spent. The questions that matter are technical ones. How do you test something that answers differently every time? What do you log, and what are you allowed to log? When does a small model do the job, and when is the large one worth the bill?
This third evening in our series is for CTOs, engineering leads and technical product people who run AI in production or are about to. After a round of introductions, we open a production AI system we built on screen and walk through what changed after deployment: what broke, what we measured, and what we would set up from day one next time. Then we hand the room over.
Small, curated and technical. 25 seats. Bring one AI feature you are running or planning to run.
Thu 21 Jan 2027 · 18:00–21:00 CET
Panenco HQ · Diestsevest 25 · 3000 Leuven
Deploying AI changes what "working" means. We look at the setup that keeps an AI feature reliable, affordable and explainable once real users depend on it.
01
A prompt or model change can make your product better or worse, and without evals you only find out from your users. We discuss how teams build test sets for non-deterministic output, where automated grading helps, and where a human still has to look.
02
When an agent calls three tools and retrieves five documents before answering, debugging starts with tracing. We look at observability for AI systems, logging inputs and outputs without creating a privacy problem, and spotting quality drift before it shows up in support tickets.
03
Prototype economics rarely survive production traffic. We cover model routing, caching, choosing between small and large models, and the point where a working feature becomes too slow or too expensive to keep.
04
Providers go down, models get deprecated, and some inputs will always produce nonsense. We discuss fallbacks, guardrails, data residency and hosting choices, and where a human stays in the loop.
Food and drinks. Programme starts at 18:30.
Name, company, and the AI feature you run or plan to run, with the part that worries you most.
A production AI system we built, on screen: the architecture, what we monitor, and the numbers on cost, latency and failure rates.
Three moments from production, each one opening a question for the room:
We collect the sharpest practices in one place and set up the next session.
25 seats. Curated audience.
We use privacy-friendly analytics and embedded content (such as YouTube videos) to improve our site. These only load once you accept. See our Privacy Policy for details.