← All articles

THE METHOD

The Tool We Recommended, Then Quietly Stopped Using

When we first tested Manus for the directory, it did what the pitch promised: handed it an open-ended, multi-step task, and it planned and executed the whole thing without hand-holding. It earned its spot in the automation category on the strength of that test, with an honest take already flagging that autonomous agents fail in harder-to-catch ways than a single chat answer.

Then we actually started relying on it for a real recurring task instead of a one-off demo. The first few runs looked fine. A few weeks in, one run quietly built a recommendation on top of an assumption that wasn't true, buried a few steps into its own reasoning where it never got surfaced back to us. Nothing crashed. Nothing errored. The output just looked confident and was wrong in a way that took real effort to catch.

That's the specific failure mode with agentic tools that doesn't show up in a quick test: the errors compound quietly instead of announcing themselves. A single bad answer from a chat assistant is easy to spot and easy to ignore. A multi-step plan built on one wrong turn near the start looks polished the whole way through, which means catching it requires reviewing the work almost as carefully as if you'd done it yourself.

We didn't pull the listing. It's still a genuinely capable tool and the honest take already warned about exactly this. What changed is how we personally use it: not as a default hand-it-off-and-walk-away tool anymore, but as something we still check closely, which quietly erases a lot of the time savings that made it appealing in the first place.

The broader rule this confirms, one we've said elsewhere on this site: an automation only earns a permanent place if you can trust the output without redoing the review work yourself. A tool that's impressive in a demo and merely fine in daily use isn't a failure exactly. It's just not the thing the demo made it look like.