00Blog
Notes from the studio.
How our apps work, how we built them, and what we learn shipping them.
2026
-
DriveStream records every drive: replay, stats and route history
DriveStream automatically logs every road trip — route, distance, moving time and top speed — with animated replay on the phone and a web companion.
-
How Rogr finds your crew offline with UWB and GPS
When there's no signal, Rogr's Find uses Ultra Wideband at close range and GPS further out, all peer-to-peer with no internet.
-
How DriveStream's fuel tracker got built
Fuel tracking was item three in DriveStream's first prompt. Here's how it went from a list item to a feature visible across the whole convoy.
-
Cribbit's one-time unlock: background listening, PiP and talk-back
Cribbit's core monitoring is free. The one-time unlock adds background audio with the screen off, Picture in Picture, talk-back and lullabies.
-
How we built DriveStream with Claude Code
From a group-chat problem to a low-poly convoy map: the prompts and decisions behind the DriveStream iOS app.
-
Rogr: Offline Walkie-Talkie for iPhone With No Account
Rogr turns two iPhones into a walkie-talkie using peer-to-peer Wi-Fi and Bluetooth. No internet, no network, no account required.
-
How we built Rogr, a walkie-talkie app, in two days
We built Rogr in two days using Claude Code, starting from Cribbit's transport layer. Here's how one prompt turned a plain communicator into a field radio.
-
How we animated Gobbl's pet in SwiftUI, with no animation tool
Gob, the pet in Gobbl's notch, is drawn in code with SwiftUI's Canvas. How its 17 moods work, who picks them, and how it stays cheap to run.
-
How we built Gobbl, a pet that lives in your Mac's notch
Gobbl went from first commit to a notarized release in a day. Why the pet is a computer, how it types along with you, and the dictation that was too slow.
-
Automated Time Tracking for Developers: No Timers, No Tagging
Manual time tracking fails because it demands effort. Here is what automated time tracking captures and how it fits a developer's actual workflow.
-
A baby monitor that needs no cloud, no account and no router
Cribbit turns two iPhones into an offline baby monitor. Here's how direct device-to-device video works, and what no cloud actually means.
-
How we built Cribbit, a peer-to-peer baby monitor
From the first prompt to TestFlight in two days — the mascot decisions, the missing encryption, and what shipping an offline app actually costs.
- One blog for every Xeve app xeve.io is now the home of Xeve, the studio behind Orma, Cribbit, Rogr, DriveStream and TidyBug. This blog covers all five, and how we build them.
-
iPhone walkie-talkie that works offline with no cell service
Rogr turns iPhones into walkie-talkies using peer-to-peer Wi-Fi and Bluetooth — the same radios as AirDrop. No internet, no account, no server.
-
Keep your road trip convoy together with DriveStream
DriveStream puts every car on one live map with speed, fuel and shared ETAs. Here's what the iOS 26 convoy app actually does.
- WakaTime Alternative: When You Need More Than Coding Time WakaTime tracks editor keystrokes. If you want full app usage, health correlations, and GitHub activity alongside coding hours, here is what to look at instead.
-
How to clean Xcode DerivedData and developer caches on macOS
Xcode DerivedData, iOS device support, npm caches and Docker VMs quietly fill your Mac's disk. Here's where each one hides and how to remove it safely.
- How to Track Coding Time in 2026: Beyond the VS Code Plugin WakaTime now tracks Claude Code sessions. Here's why that's a start, and what system-level tracking adds that heartbeat plugins alone can't give you.
- How to Set Up Automatic Time Tracking on Mac Automatic time tracking on Mac takes under five minutes to set up. Here is exactly what to install, what permissions to grant, and what you will see in the first week.
- Automatic Time Tracking: What Your Timer Isn't Telling You WakaTime's data shows the average developer codes 51 minutes a day. If you rely on manual timers, you don't know your actual number — automatic tracking does.
- VS Code Time Tracking: Automatic Coding Hours, No Timer How to track your actual coding time in VS Code automatically. Heartbeat-based tracking by project, language, and file — and why editor-only data tells half the story.
- Automated Time Tracker for Developers: No Timer Required xeve tracks every app switch, coding session, and focus block on macOS and Windows automatically — no timers, no manual input, just real data about your day.
- AI PRs Wait 4.6x Longer Because Focus Is Already Gone LinearB's 8.1M PR dataset shows AI-generated pull requests wait 4.6x longer for reviewer pickup. The bottleneck isn't the code quality. It's that reviewers have 4.2 hours of focus and it was already spoken for.
- Agent Runs Fail in the First Third. You Review the Last Third. A July 2026 study on 3,843 CLI trajectories shows failure becomes unrecoverable early — well before most developers think to intervene.
- Agents Are the Default Now. Most Developers Can't See Their Own Work. JetBrains' August 2026 survey found Claude Code overtook Copilot as the most-used AI tool at work — and that AI reshapes workflows in ways developers can't perceive.
- Kimi K3 Is $3 per Million Tokens. At That Price, Your Control Group Is Gone. Kimi K3 hit GitHub Copilot on August 6 at $3/1M tokens with a FrontierSWE score of 81.2. When frontier coding AI costs this little, selective deployment ends — and so does your ability to measure what it does.
- The Maintainer Burnout Crisis Has the Wrong Diagnosis Tidelift data shows 'loss of interest' (51%) ranks higher than 'burnout' (44%) in why open source maintainers quit. They're different conditions with different fixes.
- GitHub Agent Apps Solve One Context-Switch Problem, Not Both GitHub's Agent Apps embed Amplitude, LaunchDarkly, and PagerDuty in your PR. That solves the lookup problem. Mode shifts are something different.
- Neural Coding Didn't End Software Engineering. It Relocated It. CACM's August 2026 issue argues neural coding may be the endpoint for software engineering as a rigorous discipline. The discipline isn't ending — but where it lives has moved.
- Meetings Doubled in Two Years. They're Also Eating Your Best Hours. Hubstaff's 2026 study of 140,000 workers found meeting volume doubled and 26% of meeting time lands in the 9-11 AM window. The count isn't the problem — the timing is.
- AI Agents Didn't Just Speed You Up. They Expanded Which Languages You'll Try. A July 2026 panel study of 5,346 developers found active languages rose 2.5x at Claude Code adoption. The mechanism — an activation band of previously-too-hard technologies — reframes what AI productivity means.
- AI Tripled Commits. Releases Are Up 30%. NBER tracked 100,000+ GitHub developers across three AI tool generations. Autonomous agents add 180% more commits. Actual releases are up 30%. Here's the gap.
- The AI Governance Gap Lives at Your Desk 94% of developers use AI in the build phase. 76% of organizations have no governance process for what it ships. The gap between those two numbers is absorbed by individual developers as invisible work.
- AI Generates the Code. Nobody Writes the Rationale. ACM Queue's August issue names a new type of debt: intent debt. AI produces technically correct code while erasing the reasoning behind every implementation decision.
- 80% Use AI Coding Tools. 29% Trust What They Get Back. ACM's August myths paper buried the most important number: a 51-point gap between AI coding adoption and trust. That gap is where your productivity went.
- Coding Agents Almost Never Ask. Users Push Back Half the Time. SWE-chat analyzed 6,000 real agent sessions: agents ask for clarification in 1.4% of turns. Users push back in 44%. The mismatch is structural, not incidental.
- AI Made Juniors Faster and Seniors More Leveraged. The Bill Went to Mid-Level Engineers. LeadDev's July 2026 data names mid-level engineers as the hidden absorbers of AI productivity — doing institutional judgment work that no metric counts and no model can learn.
- You Approved the PR. Do You Still Understand the Code? Info-Tech's July 2026 study found AI code requires MORE testing despite productivity gains. The reason: approving code isn't the same as understanding it.
- Cursor Router Picks the Model. Your Session Data Loses the Signal. Cursor Router cuts AI coding costs 60% by routing automatically. The opacity means you can no longer isolate model behavior from your own performance variance.
- Shadow AI Is Already in Your Codebase Sonar's 2026 State of Code found 35% of developers use AI via personal accounts. That 42% of committed code being AI-generated? You can trace maybe two-thirds of it.
- Copilot's Code Review Skills Went GA. Writing SKILL.md Exposed a Different Gap. GitHub's AI code review agent skills went generally available last week. The teams that benefit fastest already had their review standards written down. Most engineering orgs don't.
- AI Crossed 50% of Production Code in Q2. The Measurement Frameworks Didn't. DX's Q2 2026 data: AI-authored code hit 52% of merged production output, up from 34% in Q1. Crossing the majority changes what developer metrics are actually measuring.
- AI Coding Is 22% Faster. Why Does the Next Day Feel Harder? AI tools make each session faster but create a cognitive cost that compounds across days. Here's what your session data reveals — and what to actually measure.
- Opus 5 More Than Doubled the Benchmark. Your Workflow Won't Notice. Anthropic's Opus 5 scores 43.3% on Frontier-Bench vs. Opus 4.8's 21.1%. DX Research tracked 121,000 developers and found a 65% jump in AI usage moved throughput 10%.
- Building Peaks on Weekends. Your Tracker Calls It the Same. Anthropic's Cadences report found that starting a business peaks on weekends while routine debugging drops. Most tracking tools count both sessions the same.
- The Truck Factor Assumes Someone Wrote the Code A June 2026 paper argues AI generation has invalidated the truck factor and every authorship-based knowledge metric. The number still computes. It's measuring the wrong thing.
- 31 Percent of Developer Work Is AI Overhead. Metrics Count 0 of It. A Harness study of 700 engineers found a third of the developer workday is now invisible AI overhead — review, steering, debugging AI output. No standard metric captures it.
- 94% Claim AI Gains. 67% Say the Code Needs More Testing. Info-Tech's July 2026 study caught a specific accounting error: the productivity gain is measured at code generation, not at tested and shipped.
- When Everyone Can Build, Taste Is the Moat The Solo Unicorn pitch day in NYC last week named taste, trust, and distribution as the new scarce skills. Two of those feed on one thing most solo founders don't have: good feedback loops.
- The Gap Between Tracking and Understanding Is Usually Sample Size 90 days of data looks dense. The specific events you care about — your best coding sessions, true HRV peaks, recovery crashes — may have happened 12 times. That's not a finding.
- The $3.5 Billion Admission That AI Tools Don't Know Your Workflow Microsoft and Amazon committed $3.5 billion in the same week to embed engineers inside enterprise customers. That says something about what AI tools need that software alone can't provide.
- AI's 2x Productivity Mandate Only Works on New Code A July 2026 paper tracked 802 developers achieving 2.09x PR throughput — but the gains were concentrated in newer repos and barely present in legacy ones.
- 94% of Developers Feel More Productive. The Real Gain Is 12%. A new Info-Tech survey and GitClear's 2026 Maintainability Gap report were published one week apart. Together they show how big the gap between felt and measured productivity has grown.
- Agent Loops Work Great If You Know What Actually Recurs Linear launched recurring agent workflows for teams on July 20. The hard part isn't the automation — it's knowing which tasks actually recur often enough to deserve one.
- The AI Productivity Claim That Most Organizations Can't Back Up Info-Tech's July data: enterprises with a formal AI strategy report measurable AI impact at 3x the rate. Most engineering teams are in the 20% group.
- MCP Dropped Sessions. That Changes What You Can Measure. The 2026-07-28 MCP spec removes sessions from the protocol layer. For developers measuring agentic work, something has to fill the gap.
- Everyone's Running Agents. Almost Nobody Built the Loop. Addy Osmani named it loop engineering in June. Anthropic published the guide. It's the skill after prompt engineering — designing the system that runs agents, not running them yourself.
- Copilot Has Five Models to Choose From. No Tool Tells You Which Is Right. Kimi K2.7 is now the first open-weight model in GitHub Copilot's model picker. The selection is real. The feedback mechanism for making it well doesn't exist.
- GitHub Models Closes in 13 Days. Who Actually Loses? The developers hurt by GitHub Models' July 30 shutdown aren't running production systems. They're the ones who were using it to decide.
- Linear's Agents Write 25% of Issues. The Human Work Moved. When Linear Agent converts an issue to a PR in 20 minutes, the expensive part of building software is now upstream: customer context, product judgment, intent quality. None of it shows up in a commit graph.
- The Engineers Who Try CLI Agents Aren't the Ones Who Keep Them Microsoft's first telemetry-based CLI agent study separated trial from retention. They're driven by opposite things, and the difference matters for how AI tools actually spread.
- Inline or Chat: Why Mixing AI Modes Costs More Than You Think A July 2026 field study found that combining inline AI suggestions with chat-based prompting on the same task makes both worse. Here's the mechanism.
- When Your Agent Works Overnight, Your Metrics Don't Know Claude Code shipped background subagents as default this month. When your agent commits at 2am, your time-tracker shows zero. This is the measurement model breaking.
- Grok 4.5 Was Trained on Workflows. Benchmarks Test Code. xAI trained Grok 4.5 on Cursor session data — partial edits, agent redirections, error recovery. No public benchmark measures that capability, which makes the rankings beside the point.
- Cursor Pricing Split: Why Claude Costs Extra on Every Seat Cursor's new team pricing bills Claude and GPT separately from Grok. Standard is $40/seat, Premium $120. What it costs and where the real lock-in is.
- AI Agents Helped You. They Didn't Help Your Team. Stack Overflow's 2025 survey found a 52-point gap: 69% of agent users report personal productivity gains, only 17% say it helped team collaboration.
- When You Review AI Agent PRs, 58% of the Work Isn't Code Review MSR 2026 analyzed how developers intervene in AI agent pull requests. The majority of effort is guidance and constraint enforcement, not correctness checking.
- GitHub's Infrastructure Was Built for Humans. AI Agents Don't Have Weekends. GitHub logged 88.4% availability in June as AI agents drove commits to 275 million per week. Microsoft calling AWS for help isn't the headline — the broken capacity assumptions are.
- The Credential Paradox AI Built at Entry Level PwC's 2026 report finds entry-level roles are 7x more likely to require senior judgment. That's the exact skill AI is eliminating the training path for.
- Context Fluency Is What Parallel Agent Workflows Actually Require A May 2026 paper coined 'context fluency' — externalizing tacit codebase knowledge into machine-legible form before running parallel agents. It's different from writing a spec or a prompt.
- Copilot's Bigger Return Is in Review, Not Writing Jellyfish analyzed 146,000 Jira tickets and found 2 of Copilot's 3 days saved per ticket came from review, not coding. Most teams are deploying it backward.
- Coding Agents Speed Up Projects That Haven't Used AI Before MSR 2026 tracked 400+ repos and found velocity gains from agents are front-loaded and vanish for teams already using Copilot. Cognitive complexity rises 39% regardless.
- A Quarter of Codex Requests Are Now a Full Day of Work OpenAI's June 2026 data shows 8-hour task delegations jumped from 2.1% to 25.6% in six months. When AI completes your days, the metrics measuring your week collapse.
- The People Most Replaced by AI Are the Least Worried The heaviest AI users in Anthropic's 9,700-person survey expected better outcomes on pay, job security, meaning, and autonomy — not worse.
- The Problem With Watching Yourself Work Amazon shut down KiroRank after employees gamed it with fake AI tasks. Sleep researchers call the same dynamic orthosomnia. Both are measurement problems personal analytics has to solve.
- AI's Context Window Grew 3,906x. Yours Didn't. A March 2026 arxiv paper tracked AI context windows against human attention span across nine years. The curves are moving in opposite directions.
- Microsoft's Work IQ Knows How You Work. You Don't. Work IQ reached GA on June 16, building a semantic model of your email, calendar, and meetings. Your employer's AI agents access it. You get a chatbot.
- When You Build Alone, Nobody Sees It Coming The Week named 'silent founder burnout' this week. The problem isn't overwork — it's that solo building removes every external observer who would have caught it first.
- Gartner's AI Cost Prediction Is Right. The Fix Isn't. Gartner says AI coding costs will exceed developer salaries by 2028. The number is credible. The proposed fix — governance and context engineering — gets the causation backwards.
- When a Government Pulls Your AI Stack Overnight Fable 5 launched June 9 and was suspended globally by US export order on June 12. The 72-hour window is a new data point about AI infrastructure risk that reliability metrics don't capture.
- Anthropic Writes 20% of Its Own Code. That's the Story. Anthropic disclosed Claude authors 80% of their merged code. Coverage focused on the percentage. The actual story is what the remaining 20% looks like and why it matters more.
- The Agent Shipped the Feature. What Did You Do This Week? METR's MirrorCode benchmark shows AI agents can deliver weeks of engineering work in one session. Every productivity tool we have is blind to this.
- Reviewers Like AI Code More. That's the Problem. MSR 2026's Mining Challenge found reviewers express more positive sentiment toward AI-generated code despite it having higher redundancy and cognitive complexity.
- Trusting AI Code Is the Wrong Goal The Stack Overflow survey found 84% adoption and 3% high trust. The trust framing is wrong. What you need isn't more trust — it's better failure-mode data.
- The Bill from Your AI Sprint Arrives Three Weeks Later Faros found code churn up 861% in high-AI teams. Code that passes review and ships is being removed weeks later at nearly 10x the prior rate — a second shift that appears in no standard metric.
- 30 Days to Migrate, Zero Days in Any Productivity Study Google gave developers 30 days to migrate off Gemini CLI before shutting it down Thursday. Three forced AI tool switches in 50 days. The costs never show up in productivity research.
- Kiro Makes Specs Mandatory. Mandatory Isn't Enough. Amazon's Kiro IDE won't generate code without a formal spec. Here's what the spec-first approach gets right and gets wrong.
- AI Didn't Make Us Faster. It Changed What We Were Willing to Build. Anthropic's 2026 report found 27% of AI-assisted work is work that wouldn't have been attempted at all. For solo builders, that's the more important number.
- Copilot Autopilot Is On by Default. The Developers It Helps Already Opted In. VS Code 1.124 shipped June 10 with Autopilot on by default. Behavioral data on agent expertise suggests the developers who benefit most already turned it on themselves.
- The More You Refine AI Code, the Less Secure It Gets A June 2026 paper found a 37.6% jump in critical vulnerabilities after just five rounds of AI refinement. The loop you're running to improve code is doing the opposite.
- Why Two-Thirds of AI-Generated Code Never Makes It to Production LinearB analyzed 8.1 million PRs: AI code merges at 32.7% versus 84.4% for human code. What drives the gap and what to track in your own workflow.
- WWDC 2026 Changed the Privacy Economics for Personal Data Apps Apple made Foundation Models free for small developers at WWDC. For personal analytics apps handling health data, this changes the privacy story, not the capability story.
- OpenCode Hit 172K Stars. The AI Inside Is Still Claude. OpenCode hit 172K stars in June 2026. Most of those developers still use it with Claude. What they changed is the software between them and the model.
- TypeScript Didn't Win. The AI Feedback Loop Did. TypeScript hit #1 on GitHub with 66% growth. That's partly AI's vote, not yours — and the feedback loop explains which technology is next.
- Feature Branches Are Booming. Main Branch Throughput Fell 7%. CircleCI's 28-million-workflow analysis shows AI has driven code activity up 59% while the median team's main branch throughput actually declined. More code, less software.
- The Developers Who Trust AI Agents Most Also Interrupt Them Most Anthropic's agent autonomy research found a paradox: experienced users auto-approve 40%+ of turns and interrupt 80% more often than beginners. Both rise together.
- AI Coding Output Is More Unequal Than Global Income Cursor's Spring 2026 data shows a Gini coefficient of 0.77 for AI-generated lines. P99 developers produce 46x more than the median. The average tells you almost nothing.
- Coding Agents Are Twice as Effective. Developers Still Want the Copilot. A CHI 2026 controlled study found agents complete tasks at 60% vs 25% for copilots. Then 60% of participants said they'd still choose the copilot.
- Your AI Agents Have a Dashboard. You Don't. GitHub's new Copilot app tracks every agent session in detail. It reveals something uncomfortable: your agents are better observed than you are.
- The Biometric Flow State Research Didn't Replicate A 2026 study in Empirical Software Engineering found that biometric sensors don't reliably predict developer interruptibility. The paper it tried to reproduce has been cited 68 times.
- What Copilot's Bill Shock Is Actually Telling You GitHub Copilot's metered billing went live June 1 and surfaced something bigger than the cost: developers had no idea what they were consuming.
- Refactoring Fell 60% Under AI. Nobody Tracked It. GitClear analyzed 211M changed lines: refactoring fell from 24% to 9.5% as AI adoption grew. Copy-paste now outnumbers moved code for the first time.
- AI Time Savings Have Been Flat for a Year Six quarters of DX data, 135,000 developers: AI time savings flatlined in mid-2025. Half who hit peak gains lose them the next quarter.
- 1,000 Agents Is Impressive. 95% of Your Work Doesn't Qualify. Anthropic's dynamic workflows are genuinely powerful — the Bun port proves it. But the conditions that made that port work disqualify most developer backlogs.
- The Engineering Work of Running AI Nobody's Counting Datadog's 2026 report: 69% of teams use 3+ AI models, 8.4M rate limit errors in one month. That overhead is engineering work, and it's not in any productivity metric.
- How Much Does AI Actually Improve Developer Productivity? METR surveyed 349 developers in 2026 and found AI feels 3x faster but estimates only 1.4x more valuable. Here's what to measure instead.
- The Agentic IDE Comparisons Are Measuring the Wrong Thing Google launched Antigravity 2.0 at I/O this month and the comparison posts started immediately. They're all evaluating which tool feels fastest. That's not the question.
- AI Generates the Code. Nobody Reviews It. A May 2026 study found 61% of AI-generated pull requests receive no review at all — while the same teams report rising velocity. Here's where the risk is accumulating.
- GitHub Copilot Productivity: What 2,989 Developers Found 86% satisfied, under an hour saved per week. A survey of 2,989 developers on what GitHub Copilot actually improves — and why satisfaction barely correlates with output.
- The AI Productivity Dashboard Measures the Tool, Not You GitHub's June 1 billing change ships acceptance-rate dashboards to every team. Those numbers are accurate. They're also measuring the wrong thing.
- Claude Reviews Its Past Work. Most Developers Don't. Anthropic's Dreaming feature reads past agent sessions and extracts patterns to improve future runs. It's the same problem most developers haven't solved for themselves.
- The Rules That Went Viral Are Just Code Review Advice Karpathy's CLAUDE.md four rules — think before coding, stay simple, surgical changes — are what engineering managers have said in code review for years.
- Uber Burned Its 2026 AI Budget in Four Months Uber exhausted its entire AI coding budget by April. Microsoft canceled Claude Code licenses. Neither company stopped using the tools. Here's the problem that creates.
- Copilot Had an 11-Hour Outage. Your Productivity Model Didn't Include That. GitHub Copilot logged 12 major incidents in 6 months, including an 11-hour authentication failure. When developers won't work without AI, that's a team-level production incident.
- You Don't Trust It. You Commit It Anyway. SonarSource surveyed 1,100 developers: 96% don't trust AI code, but only 48% verify it before committing. The other half ships on confidence.
- Your Company's AI Tool Mandate Is Not a Productivity Decision Microsoft told its engineers to drop Claude Code by June 30. The reason wasn't performance. That gap between enterprise mandates and individual productivity is widening.
- AI Took Junior Dev Jobs. It Also Took Their Training Ground. Stanford's AI Index shows a 20% drop in developer employment for 22-25 year olds. The displacement story is real. The apprenticeship story is worse.
- Selective AI Use Is a Performance Strategy, Not Caution Microsoft's 2026 data: the highest-performing AI users deliberately skip it more often than average. Maximum AI use doesn't maximize output.
- The Open Source Maintainers Paying for Your AI Productivity curl shut down its bug bounty in January. tldraw stopped taking external PRs. Jazzband sunsetted. Your AI output has a cost — it just lands somewhere else.
- Engineering Leaders Know Their Metrics Are Broken. They're Using Them Anyway. Harness surveyed 700 engineering leaders this month. 89% trust their metrics. 94% say key factors are missing from those same metrics. Both cannot be true.
- The Slot Machine in Your Terminal Karpathy runs agents 16 hours a day. Ronacher barely sleeps. This looks like enthusiasm but has the structure of a slot machine.
- Deployment Frequency Is Up. Your Change Failure Rate Too. AI boosted deployment frequency across engineering teams — but Cortex's 2026 benchmark found change failure rates rose 30% in parallel. DORA's core assumption is breaking.
- You Can't Delegate What You Can't Specify Anthropic's 2026 report found developers use AI in 60% of their work but fully delegate 0-20% of tasks. That gap is a specification problem, not a model problem.
- Git Pushes Are Up 78%. Most of Them Aren't Human. Microsoft's May 2026 AI diffusion report cites surging git activity as a productivity win. GitHub is processing 275 million agent commits per week. Those are different things.
- The Metric AI Productivity Research Keeps Ignoring Two years of IDE logs from 800 developers found AI users delete code 13x more per month. That deletion is review labor — and it shows up nowhere in standard metrics.
- The Control Group Refused. That's METR's Most Interesting Result. METR abandoned their AI developer productivity RCT because 30-50% of developers wouldn't work without AI. The measurement breakdown is the finding.
- AI Doubled Your PR Count. Review Didn't Scale. Teams using AI are merging 98% more pull requests but not shipping twice as fast. Faros.ai's 2-year study of 22,000 developers shows where the gains disappeared.
- Tokenmaxxing Is the New Commit Count Jellyfish tracked 7,548 engineers in Q1 2026 and found developers burning the most tokens produced 2x the output at 10x the cost. Volume is not value.
- Copilot Moves to Usage Billing. Now Prove the ROI. GitHub Copilot switches to token-based billing on June 1. Most teams have no data to answer the question the billing change is now asking.
- Vibe Coding Fixed the Wrong Bottleneck Moltbook was breached three days after launch. Lovable exposed projects for 76 days. These aren't anomalies — they're what happens when you ship code you don't understand.
- AI Won't Hollow Out Your Skills. Your Behavior Will. Anthropic's study found AI reduces comprehension 17% on average. That average hides a 65% vs 40% split determined entirely by how you use the tools.
- AI Sped You Up 20%. It Actually Slowed You Down 19%. A METR study found experienced developers were 19% slower with AI tools — yet believed they were 20% faster. The gap is real and you cannot feel it.
- You Code for 52 Minutes a Day. AI Optimized That. 93% of developers use AI coding tools. Productivity gains are stuck at ~10%. The reason is structural: actual code writing is a tiny slice of the job.
- Four Agents at Once Will Wipe You Out by 11 AM Running parallel AI coding agents is genuinely faster — Simon Willison confirmed it. He also said he's mentally exhausted before noon. Here's what that trade-off looks like in the data.
- The Four-Day Week Trial Exposed Five Hours of Hidden Waste The biggest 4-day work week study found workers match output in 33 hours vs 38. That 5-hour gap exists in your week too — tracking reveals where it hides.
- Solo Founders Don't Have a Speed Problem The solo unicorn narrative is landing — Medvi, $401M, two people. But what actually made it work has nothing to do with coding faster.
- AI Saved You an Hour. Researchers Tracked Where It Went. Two 2026 studies found AI tools don't reduce developer workload — they expand it. The time savings become more tickets, more scope, and higher burnout.
- Going Async Didn't Give You Your Focus Time Back Remote work promised developers more deep work by cutting meetings. Five years later, the average focus session is 13 minutes and still falling.
- AI Makes You Feel Faster. Controlled Studies Disagree. Two recent studies show developers systematically misjudge AI's impact on their output. The data tells a different story — and it matters for how you measure your work.
- Context Switching Is Killing Your Productivity — Here Is the Data The average developer switches apps 300+ times per day. Research shows each switch costs 23 minutes of focus. Here is how to measure and reduce the damage.
- How to Measure Developer Productivity Without Surveillance Developer productivity metrics that actually work — without screenshots, keyloggers, or invasive monitoring. Track focus time, context switches, and coding output automatically.
- How to Track Coding Time Across Multiple Projects Automatically Stop guessing where your coding hours go. Automatic tracking across VS Code, Xcode, terminal, and Claude Code — per project, per language, no manual timers.
- I Used an Open-Source SEO Plugin for Claude Code — It Found 25 Issues and Fixed Them in One Session How I used claude-seo, an open-source skill package for Claude Code, to run a full technical SEO audit across 9 categories, score my site, and auto-fix 15 issues in a single conversation.
- I Built an iOS App With Claude Code — HealthKit, Widgets, BLE, and TestFlight A step-by-step walkthrough of building the xeve iOS companion app entirely through Claude Code prompts — SwiftUI, HealthKit, CoreLocation, CoreBluetooth, WidgetKit, SwiftData offline sync, and TestFlight deployment. Every prompt, issue, and workaround documented.
- I Built a Native macOS App With Claude Code and Never Opened Xcode A complete walkthrough of building a production macOS menu bar app — SwiftUI, Supabase, BLE heart rate, Sparkle auto-updates, code signing, and notarization — entirely through prompts in Claude Code. Every prompt, every issue, every fix.
- Does Music Affect Coding Productivity? What My Data Shows I tracked my Spotify listening history alongside coding sessions for 3 months. Here is what the data reveals about music genres, focus, and code output.
- We Built an Energy Forecast That Predicts Your Best Hours — Like a Weather Report for Productivity xeve now predicts your hourly energy levels 7 days forward using historical work patterns, sleep, recovery, and meeting load. A heatmap calendar shows when to schedule deep work, and a burnout monitor warns you before you crash.
- How Much Time Do I Actually Spend Coding? Most developers overestimate their coding time by 2-3x. Here is how to measure it accurately and what the data typically reveals.
- The Exact Prompts That Took Our UI From Generic to Distinctive — A Real Design Journey With Claude Code We went through glassmorphism, Three.js particles, and Gyroscope clones before finding the Teenage Engineering aesthetic. Here are the actual prompts, the mistakes, the corrections, and the 7-step framework that emerged.
- Quantified Self for Developers: Coding, Sleep, and Focus Tracking WakaTime tracks code. Whoop tracks sleep. Neither explains your best days. How to connect all four data layers and what the correlations reveal.
- How to Track Screen Time on Mac (Beyond Apple Screen Time) Apple Screen Time is basic. Here are the tools that give developers real insights into how they spend time on their Mac — with coding analytics, categories, and trends.
- WakaTime vs RescueTime: Which Should You Use in 2026? A detailed comparison of WakaTime and RescueTime — what each tracks, pricing, pros and cons, and why neither gives you the full picture of developer productivity.
- The Mistakes Claude Code Makes — And What They Teach You About AI-Assisted Development We built xeve with Claude Code. It shipped light mode across 155 files, built an MCP server in an hour, and published to npm. It also broke the build three times, forgot half the fix, and referenced UI that did not exist. Here is what we learned.
- AI Weekly Digest — Cross-Platform Pattern Detection You Would Miss Every week, xeve feeds your aggregated data to an LLM and gets back personalized insights: productivity patterns, health correlations, anomalies, and actionable recommendations.
- How to Run a 155-File Migration With AI Agents Without Losing Your Mind Practical guide to large-scale codebase migrations with Claude Code — parallel agents, sed scripts, build-after-every-phase, and the specific failure modes to watch for.
- Calendar Sync — Measuring Meetings vs. Maker Time xeve now syncs your Google Calendar and shows exactly how much of your week is meetings vs. deep work. Meeting hours, average duration, busiest days, and recurring meeting analysis.
- Communication Tracking — How Much of Your Day Is Slack, Email, and Calls xeve now tracks time in communication apps with per-channel and per-contact breakdowns. See your daily communication patterns, peak hours, and which conversations consume the most time.
- The Energy Score — A Single Number for How Ready You Are to Work xeve now computes a daily Energy Score from 0-100 based on sleep, activity, heart rate, focus quality, and screen time balance. One number that tells you if today is a push day or a recovery day.
- Goals — Set Daily Targets and Track What Matters xeve now supports personal goals. Set minimum coding time, maximum screen time, step targets, and focus thresholds. Track daily progress against your own benchmarks.
- The iOS Companion App — HealthKit, Location Tracking, and Widgets xeve for iOS brings health data, location awareness, and glanceable widgets to your personal analytics. HealthKit steps, sleep, heart rate — plus home/work detection and three widget types.
- Light Mode — A Warm Industrial Palette That Actually Feels Right xeve now has a light mode. Not clinical white — warm linen backgrounds, bold contrast, and the same orange accent. Built with CSS custom properties and zero new dependencies.
- Ask Claude About Your Day — xeve Now Has an MCP Server Connect xeve to Claude Desktop or Claude Code via MCP and query your productivity, coding, health, music, and GitHub data conversationally. Open source, 9 tools, zero config.
- Project Breakdown — See Exactly Where Your Coding Time Goes xeve now breaks down coding time by project. See which codebases consume the most hours, track daily project allocation, and understand how your engineering effort distributes across repos.
- One Day, Five Bugs, Three Features — What Shipping Fast Actually Looks Like We shipped light mode, an MCP server, and fixed invisible bugs that had been silently breaking auto-updates for three releases. A raw look at maintaining a multi-platform product.
- Every Song, Every Album — How xeve Tracks Your Spotify Listening History xeve polls the Spotify API every 90 seconds to capture your full listening history — track name, artist, album, album art, and duration. See your music habits alongside productivity data.
- The Timeline — See Your Entire Day as a Gantt Chart xeve now shows your daily app usage as a horizontal timeline. Every app switch, every session, color-coded by category. See your day at a glance and spot patterns in how you work.
- Website Tracking — Where Your Browser Time Actually Goes xeve now breaks down your browser time by individual website. See which sites are productive, which are distractions, and how your browsing patterns change across the week.
- The Weekly Report — This Week vs. Last Week, Every Metric xeve now generates automatic weekly comparison reports. Screen time, coding, communication, music, GitHub, health — all compared week-over-week with percentage deltas.
- The Correlation Engine — 19 Metric Pairs That Reveal What Actually Drives Your Productivity xeve auto-computes Pearson correlations across 19 pairs of daily metrics. Sleep vs. coding. Steps vs. focus. Music vs. app switching. Each with a plain-English interpretation.
- Export All Your Data — CSV or JSON, Any Date Range Your data belongs to you. xeve now supports full data export in CSV and JSON for every data type — app sessions, coding time, health, music, GitHub, locations — with date range filtering.
- GitHub Activity Sync — Commits, PRs, and Reviews in Your Dashboard Connect GitHub and see your commit history, pull requests, and code reviews alongside your other productivity data. Daily contribution charts, language breakdown, and cross-repo activity.
- GitHub Member Sync, AI-Enriched Profiles & Smart Project Filtering xeve now imports your entire GitHub org — members, contributors, and activity. AI infers roles from commit patterns, generates project descriptions from READMEs, and the People page shows everyone in your org, not just xeve users.
- Sync AI Meeting Summaries into Your Org Dashboard Connect your meeting analyzer to xeve and see AI-generated summaries, key decisions, action items, and blockers — all organized by date in your org dashboard.
- Auto-Extract Tasks from Meeting Notes — Assign, Track, Complete xeve now parses action items from your meeting summaries, auto-assigns them to team members by name matching, and lets everyone track completion in real time.
- VS Code Extension — Automatic Coding Time Tracking Without WakaTime xeve tracks your coding time in VS Code via heartbeats. Language, project, file activity — all captured automatically and synced to your dashboard. No API keys. No config files.
- xeve for Windows — A Native WinUI 3 Tracker, Not Electron We built a native Windows app using WinUI 3 and .NET 8. System tray tracking, Win32 foreground detection, project extraction from window titles, and the same design system.
- Real-Time Heart Rate from Any BLE Monitor — Whoop, Polar, Garmin xeve connects to Bluetooth Low Energy heart rate monitors on both macOS and iOS. See your heart rate in real time, log it alongside your work, and discover how stress affects your productivity.
- Introducing xeve Enterprise — Role-Based Dashboards with AI Insights Deploy xeve across your company. Each executive gets an AI-powered dashboard tailored to their role — not generic charts, but actionable intelligence about team performance.
- Track Your Plex Watch History Alongside Your Productivity Connect Plex to xeve and see how your media habits fit into your daily routine. Watch history, now playing, and media-productivity patterns — all in one dashboard.
- Streak Tracking and Heatmaps — Visualize Your Consistency GitHub has contribution heatmaps. xeve has them for everything. See your daily consistency across productivity, coding, exercise, and music — with streak tracking to keep you accountable.
- How Focus Tracking and Distraction Nudges Changed My Work Habits xeve now tracks deep work sessions, measures context switching, and nudges you when you drift to unproductive apps. Here is how focus tracking works and what I learned.
- Introducing Teams — Shared Productivity Dashboards xeve now supports teams. Create a team, share an invite link, and see aggregated productivity metrics across your group — without compromising individual privacy.
- How I Track Every App I Use on macOS (Automatically) Manual time tracking never works. Here is how automatic app tracking captures every app switch, window title, and category — without you lifting a finger.
- Does Sleep Affect Coding Productivity? My Data Says Yes I correlated 3 months of HealthKit sleep data with my daily coding output. The results were not surprising — but the magnitude was.
- Track Your Claude Code Sessions with One Command AI-assisted coding is the new normal, but most developers have no idea how much time they spend in Claude Code. Here is how to track it automatically.
- Building a Personal Analytics Platform with Supabase and Next.js The architecture decisions behind xeve — why Supabase over Firebase, why native Swift over Electron, and what I learned building a full-stack analytics platform as a solo developer.