What a computer-using agent actually does all day
Two weeks of screen recordings, one honest tally: where the agent earned its keep and where it quietly wasted an afternoon.
Two weeks of screen recordings, one honest tally: where the agent earned its keep and where it quietly wasted an afternoon.
Why the people who delegate well to juniors delegate well to models, and what that means for hiring.
Everyone agrees you need evals. Nobody agrees whose job it is. We work through a real one, badly, then better.
Running a second model over the first model's output caught real mistakes. It also doubled the bill and added a step nobody owned.
Most AI advice is written by people who never have to ship anything. This is the other version: one week of real work, reported back.
A new episode every week, wherever you already listen.