I promised the second part would be about the difference between a website, an app and software. That part is still coming.
I promised the second part would be about the difference between a website, an app and software. That part is still coming. But over the weekend of 1 and 2 August a bug was found in the Ptáček app, and the promised difference played out live instead of in a table.
Showing it on a real bug is more honest than explaining it with examples out of my head. So the order changes.
There's a second reason. I want you to see what development with AI really looks like. Not just the finished demo at the end, but also the moments when it breaks, when you look for the cause and when you step back twice before it works. Vibe coding is sold as magic. In reality it's a craft with bugs that get fixed.

What broke and why silence is worse than a crash
Every five minutes Ptáček asked the system calendar whether anything was coming up. The problem was how it asked. For every query it opened a new connection and didn't close the old one. After roughly four hours of running, macOS refused to give it another connection.

The second half of the bug was worse. Instead of an error message the app got an empty response and treated it as valid. No calendars, no meetings. It even deleted the fly-bys it already had scheduled. Meanwhile the settings showed green that calendar access was working.
For an app that is supposed to fly across the screen a few minutes before a meeting, this is especially unpleasant. A person relies on it and stops watching the calendar themselves. When it then stays silent instead of saying I don't know, it's worse than if it hadn't existed at all.
For a reminder app the worst error is not a crash. The worst is silence.

Who found it: two audits and one correction
I put two independent audits on the cause. Two different models, each with its own brief, both with read access to the code and to the running app. They didn't know about each other.
The first found the cause directly in the code: three places where a new connection is opened instead of reusing the old one. It immediately ruled out fixes of the kind “let's raise the connection limit”, because that would only postpone the same problem.

The second confirmed it with its own finding from the log and on one point corrected the first.
The takeaway I draw from it: one audit gives you an unconfirmed conclusion. Two independent audits give agreement on the cause plus one correction that the first wouldn't have come up with alone.
No user ever met the bug. It was caught by our test process on Saturday evening, not by a stranger on Monday morning.

What was fixed: architecture, not a patch
The fix didn't mean “let's raise the limit” or “let's add a retry”. What changed is how the app talks to the calendar.
One connection for the whole run of the app instead of a new one for every query. A bug never looks like emptiness again. When the app doesn't know, it says I don't know and lets the last confirmed plan stand, instead of translating emptiness as “no meetings”. Reaction to a changed meeting within a second instead of up to five minutes.
There is a difference between “it doesn't show red any more” and “it doesn't happen any more”. The second means understanding why it happened and rewriting that part, not covering it up.

A map of the weekend: development is not a straight line
Looking back, it seems a clean story. A bug, two audits, a fix, done. In reality it didn't run so straight. Over the weekend it was nineteen milestones, three bugs along the way and three steps back, when a running attempt was thrown away and we started again from the last working state.
Demo videos never show this.

Independent check and the numbers: from 7.5 to 9.2
After the fix came a second wave of checks, independent of both previous audits.
Ten million simulated calendar steps without breaking a single rule. A million operations with a single live connection. Exactly what was missing on Friday. Fifty-nine tests directly in the code and on top of that sixty-seven tests from our own QA lab. Forty hand-designed scenarios in eight categories.

That lab can even recognise a planted faulty version. It is given a deliberately broken model and has to shout NO GO on its own. If it didn't shout, it's useless. Tests that never find anything are decoration, not control.
Eighty to twenty and three days
The app was built eighty percent by AI, and for the remaining twenty percent a person steered and checked it. The brief, the risk decision and the last word always belonged to a person. Without that control the weekend would have gone differently, because an audit that never says NO GO on its own is useless.
The point isn't that bugs disappear. The time shortens between a bug arising and its being fixed and documented.

Why this happens even to Apple
The library from Apple that Ptáček uses for reading the calendar was used differently from how Apple intended. The documentation says so in one sentence hidden in a thousand pages.
It happens to Apple and Google too, companies with decades of processes behind them. The difference isn't whether a bug happens, but how quickly you find and fix it.

Three questions for every supplier
How many tests ran and what failed in them? If someone tells you nothing ever failed, it wasn't tested.
What keeps running? Whoever has no map of further testing hasn't done any.
What scale is it? A website, an app and a connection to a twenty-year-old company system are three different leagues. The risk grows the same way the price does.
The form of a prompt is always perfect. The substance isn't.

AI is a turbo, not an autopilot. And a turbo without a driver ends up in the ditch as fast as it drives into it.
The article was first published on LinkedIn. Original article on LinkedIn

