OpenAI has added a voice to the ChatGPT desktop app that lets you steer several agents at once. You start a task, interrupt it, rearrange priorities, and all of it by mouth.
OpenAI has added a voice to the ChatGPT desktop app that lets you steer several agents at once. You start a task, interrupt it, rearrange priorities, and all of it by mouth.
I've been putting this together by hand for a few months now, and the voice is the least interesting part of it. The interesting part is the question that follows: when everything can be built this fast, who can tell it's wrong? I'll show you one day and the answer at the end.
Morning: twenty minutes of walking
The best brief of the day is created outside. Not at the keyboard.
The day doesn't start at the computer. It starts outdoors, with an earbud in. I dictate what I want to do that day, why and what I'm afraid of. I don't phrase it, I just get it out of my head. In twenty minutes of walking that gives more context than I'd write in an hour, and better, because I speak in whole thoughts.
The transcript makes mistakes on names and numbers, I admit. But for thinking out loud it works better than anything I've tried.

Late morning: eight terminals side by side
Eight runs at once. Not because I'm fast, but because I don't wait.
At the desk it already works differently. For years I've been tuning my own environment: eight terminals side by side, a grid that builds itself, over two thousand lines of my own configuration and helpers that work at hand. What I was missing, I wrote myself.
The practical consequence: in one late morning, five things run at once. A brief for a tender for a logistics company, a document for a partner, an update of my own product, plus content and the internal systems the company stands on. I don't wait for one to finish before I start the next. That isn't a trick, it's capacity. One person gets through what used to need five specialists.

The second half of the late morning is boring and the client never sees it. I go into every call having read everything that has been produced for it. For the last partner negotiation that was a hundred and nineteen files, from the transcript of the first call to a test scan of his product. Thanks to that I don't speak in generalities, but to a concrete sentence the other side said three weeks ago. That's the difference between a meeting that produces a decision and a meeting after which a summary gets sent.
On top of that I'm now trying Orca, an app that steers those terminals as one whole. The panels finally know about each other and the context between them isn't carried by a person. I've had it only a short time, so no big conclusions. When I have weeks of miles on it, I'll write how it turned out and what I don't like about it.
Afternoon: what actually got made in the day
Not how many hours. What came out of them.
This is where you can see what this way of working really makes possible. And exactly here begins the problem I'm writing the article about.

Three things from the last few weeks, not from the whole year.
An app in half a day. In the summer I built a small Mac app that watches over meetings. Overnight four versions, tests from eighteen to fifty-two, two independent audits, a rating of 9.2 out of 10. Free and with the source code, as a public beta. Now there are several hundred downloads behind it and I'm preparing it for Reddit and other communities, where more people than me will break it. I'll keep writing about what comes out of that. The whole story of that first night is in the article “Half a day for the app, a day and a half to make it fit for strangers”.
A supplier measured. Before we signed a collaboration with one partner, I ran his service through 35 real transcripts from our calls. 35 out of 35 without an error and 9.7 times cheaper than our existing model. They themselves advertised 9.1 times. I came to the call with numbers, not impressions. The whole measurement with the methodology is in the article “9.7x cheaper than Claude Sonnet”.

Hundreds of articles for one client. A publishing system that writes in her voice, and she approves from her phone. An article for a few euros, not thousands. How it is all built I broke down in the article “Your website can work while you sleep”.
But I like most the numbers that I don't say. One client, a marketing agency, calculated that the system saves him forty hours a month. I didn't calculate that number, it came from him. And that's exactly the kind that interests me, because it can be checked in his timesheet, not in my presentation.
This isn't about typing speed. It's about not waiting, several things running at once and every output being measured. Behind that speed are fifteen years of projects and over five thousand hours with AI, otherwise nothing usable would come out.
The first version is always decent and useless. I threw this article away three times and wrote it again, because something wasn't right. The speed isn't in the first attempt coming out well. It's in my recognising it and throwing it away the same day.

And exactly here is the question from the introduction. When everything can be built in a day, who can tell it's wrong? In my case the answer is boring: fifteen years of the same questions in different industries. I can tell in seconds what will fall over on Friday evening, but not because I'm smarter than the model. Because I've seen it fall over before. With a glasses hinge that mustn't pull hair. With a warehouse that mustn't lie to accounting. With a payment gateway where a mistake is counted in transactions. The model generates a proposal in minutes, but never says on its own I don't know. That decision stays with a person, and it can be made quickly only when they have put in the practice behind it.
Afternoon, second half: the opponent
A mistake that shouts is good news. The one that stays silent is a disaster.

And now the part almost every supplier leaves out. That is exactly where the losses happen.
But after six hours on one thing, even the most experienced person overlooks a blind spot. That's why everything created that day goes through a second model as well. Not the same one, a different one. It gets its own brief and looks for holes. Because one model doesn't argue with its own work. It praises it. Always.
That isn't my impression. Zapier had it independently measured how models do on 657 real business tasks, and whether they follow the given rules was assessed too. The best fell below 50 percent. Translated: a top model does a real business task correctly and by the rules in roughly half the cases. Without a second check nobody sees the other half.

Besides the second model, everything also goes through my own QA lab. Sixteen rules, a verdict to release or not, with no excuses that this time it's good enough. The website is also run automatically at five widths before deployment, not by eye. Then a manual click-through and a check of production. Last time it found a link leading to a place that didn't exist on the page.
With that meetings app it looked like this. The second model didn't return one finding. It returned four conditions without which I shouldn't have released it. Two are worth describing.
The uninstaller deleted by a name mask. On my computer that was fine. But the mask could also hit somebody else's app. The fix was to verify the identifier, not the name.
When loading the calendar failed, the function didn't return an error. It returned an empty list. The app pretended you had no meetings and happily showed green. That isn't a crash you'd notice. That's silence.
The worst error is not a crash. The worst is silence, when the system claims everything is fine. And that holds for every system anyone sells you. Ask who argued against it.

Evening: the write-up
The last thing of the day is the most boring and almost everybody gives it up. Everything that was decided and created that day gets written down. Not in my head, in the system. From those notes I built a digital twin: it knows my decision rules and tells me where I contradict myself. Sometimes sooner than a person would. The next morning I don't start from zero and neither does the model, because it knows the context of all the projects, not an average from the internet.
Without that, every setup is a dead folder in three months. With it, things add up.
Do you have someone by your side who tells your AI it's wrong?
What to take from this if you're not a developer
The idea isn't technical, even if the article sounds that way. Jarvis isn't one clever assistant. It's three layers: a fast input, a memory so that nothing repeats, and someone who says it's wrong before the client sees it. The first two anyone can get today in an afternoon. The third one decides.

And almost everybody leaves out the third. That's how you get thirty-page documents nobody read, and systems that show green even when they're silent. When you choose an AI supplier, this is the question to ask: who at your place argues against what you build?
We build these three layers for companies outside development too. If you're curious what such a day would look like at your place, write to me, or just take a 15-minute intro call: cal.com/transformuj.ai/30min
The article was first published on LinkedIn. Original article on LinkedIn

