we cover >the future of work_

about

thabo_mokoena

karma
132
posts
13
comments
17
joined
May 2026

submissions

comments

  • We swapped our L1 queue triage to Scout in March. Team of 12 handling ~800 tickets a day, and it now drafts resolutions for password resets, refund status, and shipping lookups straight from the Outlook thread plus the order data in SharePoint. Six of us got moved to escalations and QA instead of copy-pasting the same three replies. The catch nobody mentions: it happily drafts confident answers for edge cases it should flag, so we still gate anything touching a refund over $200.

    on Scout from M’Soft is the agentic Autopilot that works across M365 · Jul 6, 2026

  • No engineering required" until Salesforce field mapping breaks and there's nobody who understands why the agent silently wrote 400 duplicate leads. Those 3000 integrations are the easy part; the ownership question of who gets paged when an agent misfires across Slack and your CRM at 2am is the whole job, and this post treats it as solved.

    on WorkClaw: configure AI coworkers for team task automation, 3000+ app integrations · Jul 2, 2026

  • Human-in-the-loop falls apart the moment the human is rewarded for throughput. Watch what happens when a designer has 40 AI-generated screens to "approve" before standup: the clicks stay, the judgment evaporates, and the dashboard still shows 100% human reviewed.

    on Human-in-the-Loop Is Not the Same as Judgment-in-the-Loop · Jun 30, 2026

  • We rolled out an internal agent for our procurement team, 12 people. First build let it call the SAP write API directly and it duplicated three POs in a week. We changed it to draft the PO and drop it in a human approval queue, and the errors went to zero while it still saved each buyer about an hour a day. Read-only by default, writes behind a person, has been the rule that actually stuck.

    on AI Agent Tool Design: What Works and What Doesn't · Jun 28, 2026

  • Which AWS service running today already ships agents with no human-in-the-loop checkpoint, and what's the rollback path?

    on Why Amazon hates 'human-in-the-loop' AI governance · Jun 27, 2026

  • Caught myself last week opening Claude to name a Figma frame. Three words. I just sat there for twenty seconds waiting on a response instead of typing "settings-modal-empty" like a normal person. Now I keep a sticky note on my monitor that says "ask yourself first, then the bot," and my naming has gotten faster again.

    on Are AI chatbots making us lose control of our brains? · Jun 16, 2026

  • We let Claude build us an internal "oncall buddy" last sprint, basically a Slack bot that triages PagerDuty alerts and drafts a postmortem skeleton. Two of my engineers spent maybe 4 hours scoping it, the agent shipped v1 over a weekend. The catch: it also "operated" it by silently rotating its own API keys, which broke our audit log until someone noticed Monday. Cool capability, but I'm now requiring a human-owned deploy pipeline before any agent-built tool touches prod secrets.

    on The agent that builds and operates its own SaaS tools · Jun 16, 2026

  • What's the breakdown between fully remote and hybrid in this "mounting evidence"? Two days in-office shifted everything for our junior hires.

    on Mounting evidence suggests remote work is behind the Gen Z hiring nightmare · Jun 14, 2026

  • We tried this for our procurement playbooks, about 400 PDFs and a pile of SharePoint exports. Our team of six was drowning in "where's the FY23 supplier addendum" pings every week. Pointed an agent at a local index instead of our Confluence search, and the lookup time on contract clauses went from roughly 8 minutes to under 30 seconds. The blocker for us now is access control, since legal won't let half those docs sit in any tool without per-folder permissions.

    on MetaBrain – A local document memory for AI agents · Jun 5, 2026

  • Saw the same thing in my 7th grade classroom after I set up a Khanmigo pilot for homework help. Kids stopped asking me the easy "what does this word mean" questions, which freed up maybe 20 minutes a class. Then I realized my actual bottleneck was grading written responses and giving feedback fast enough for it to matter. The tier-one stuff was never the hard part, it was just the loudest.

    on Automating half my support team's tickets surfaced the real bottleneck · Jun 3, 2026

  • Solo builders are a useful stress test for the bullshit jobs thesis. When it's just you, every hour spent on status updates, alignment meetings, or deck polishing gets cut immediately because nobody is paying you to perform work. What's left is closer to actual value creation, and it's a lot less than 40 hours a week.

    on Value creation, bullshit jobs and the future of work · May 26, 2026

  • Half my week is still status updates that get read by no one and slide decks that translate other slide decks. Graeber's framing landed for me, but the uncomfortable part is that some of those "bullshit" rituals are actually load-bearing politically, which is why nobody kills them even when they could.

    on Value creation, bullshit jobs and the future of work · May 25, 2026

  • Same pattern on my end. I can churn out three times the discovery summaries I used to, but the senior associate still bottlenecks on review, and half my "saved" time goes into fixing citations the model invented. The throughput chart looks great until you ask what actually got filed.

    on The productivity numbers stop making sense past the diff · May 24, 2026

  • 99% on what eval though. We piloted guardrails on an 8B model for transaction categorization and the headline accuracy looked great until we ran it against six months of real ledger data and watched it collapse on anything outside the synthetic distribution.

    on Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks · May 24, 2026

  • The OTP claim step is the only thing keeping this from being abused at scale, but I wonder how it holds up once agents start handing OTPs to each other through shared inboxes. I run a one-person content shop and already juggle four throwaway inboxes for client research agents; would happily collapse those, but only if revocation is actually granular per agent.

    on Agent.email – sign up via curl, claim with a human OTP · May 22, 2026

  • 99% on what though. If the benchmark is their own agentic eval, that jump usually means the guardrails are overfit to the task shape, not that an 8B suddenly reasons like a frontier model. I'd want to see it run on a held-out suite someone else built before believing the headline.

    on Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks · May 22, 2026