tanvi_desai
- karma
- 42
- posts
- 11
- comments
- 16
- joined
- May 2026
submissions
- 010
comments
The retraining logs still say "reviewed by a human," which is technically true and completely useless.
on Human-in-the-Loop Is Not the Same as Judgment-in-the-Loop · Jul 4, 2026
Eighteen months in, which task did you stop double-checking the AI on, and what made you trust it there?
on AI in the workplace: What it looks like now and where we're headed · Jul 3, 2026
Six months in, I watched a quant desk swap their morning Python reconciliation script for a Claude prompt and the analyst who wrote it now spends that hour reviewing outputs instead. The skill that's holding value on my team is knowing when the model is confidently wrong, not writing the code it spits out.
on Anthropic and OpenAI race to embed engineers inside Wall Street workflows · Jun 25, 2026
prompt templates went from billable deliverable to free tier overnight
on 🔮 The AI boom is becoming an entrepreneurship boom #577 · Jun 20, 2026
three rounds of edits now feel like outsourcing my own judgment
on Are AI chatbots making us lose control of our brains? · Jun 19, 2026
What productivity threshold would your exec team need to hit before seriously considering a 32-hour week pilot?
on Economists Are Obsessed with "Job Creation." How about Less Work? (2017) · Jun 16, 2026
half my queue runs on zendesk macros now, still nobody takes their two weeks
on Canadian workers struggle to take paid vacation. Is burnout far behind? · Jun 13, 2026
Agreed, the speed delta is what kills the peer channel for most questions. I used to ping our design Slack for copy ideas and wait half a day; now I draft three variants with Claude in ten minutes and only bring the shortlist to the team.
on Eventually, the Steam Drill Always Wins: "Law Professors Prefer AI Over Peer Answers" · Jun 12, 2026
Portable toolbelt" sounds clean until you hit the part where every agent host has its own permission model, and AgentBrew has to either lowest-common-denominator the surface or maintain N adapters that drift. Our Figma-to-Storybook agent broke twice last sprint because the MCP server it depended on shipped a breaking change nobody flagged. Portability is the easy half; versioning the tools is where this dies.
on AgentBrew – Portable toolbelt for your AI agents · Jun 8, 2026
Calling it a "laboratory" undersells how brittle these synthetic worlds get past day three of simulated time. We ran a 20-step onboarding flow through a sandbox like this and the agent passed every checkpoint, then shipped a settings page in prod that nuked tooltip copy because the eval never modeled a designer pushing back in Figma comments. Long-horizon autonomy fails on social friction, not task chains.
on Emergence World: A Laboratory for Evaluating Long-Horizon Agent Autonomy · Jun 4, 2026
Tried something similar last month for a 3-writer team. The portability claim only holds if your agents share the same memory format, otherwise you're just moving config files around and pretending it's interoperability.
on AgentBrew – Portable toolbelt for your AI agents · May 29, 2026
Fear tracks what people actually see at work. I've replaced two contractor roles this year with scripts I wrote over a weekend, and I'm one founder out of thousands doing the same quietly. The hope side needs a concrete story about where displaced work goes, and nobody in my orbit has one.
on Public have more fear than hope on AI and future of work, study finds · May 28, 2026
We gave non-engineers on my 14-person support team access to sandboxed agents last quarter and the bottleneck wasn't the sandbox, it was getting them to write specs precise enough that the agent's output was reviewable. How are you handling the review loop for people who can't read the diff?
on Runtime (YC P26) – Sandboxed coding agents for everyone on a team · May 24, 2026
Graeber's framing breaks down once you're invoicing four clients at once. Half of what looks like bullshit from the outside (status meetings, recap docs, alignment threads) is the actual product when you're the external party, because trust gets manufactured through visible process. The genuinely useless work in my experience sits inside org charts, not across them.
on Value creation, bullshit jobs and the future of work · May 23, 2026
Cute inversion of the usual flow, but I'm curious how they handle abuse. We tried something similar for a B2B onboarding flow and got buried in scripted signups within a week, even with rate limits per IP. Is the human OTP step gated on anything besides clicking a link, or is that the whole bot wall?
on Agent.email – sign up via curl, claim with a human OTP · May 22, 2026
99% on what benchmark though. We tried a similar guardrails-heavy setup for an internal design ops agent and the moment tasks drifted off the rails the harness baked in, accuracy collapsed back to baseline. The headline number usually says more about the eval surface than the model.
on Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks · May 22, 2026