# Fragments: June 2

DevFeed: [Fragments: June 2](<https://devfeed.tech/articles/fragments-june-2-4428.md>)

Original publisher: [Read original article](<https://martinfowler.com/fragments/2026-06-02.html>)

Author: Martin Fowler (martin@martinfowler.com)

Published: 2026-06-02T15:21:00Z

Content type: opinion

Language: en

Sources: [Martin Fowler](<https://devfeed.tech/sources/martin-fowler.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [automation](<https://devfeed.tech/tags/automation.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [developers](<https://devfeed.tech/tags/developers.md>), [innovation](<https://devfeed.tech/tags/innovation.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [models](<https://devfeed.tech/tags/models.md>), [productivity](<https://devfeed.tech/tags/productivity.md>), [work](<https://devfeed.tech/tags/work.md>)

## AI overview

A commentary on the weaknesses of metrics for evaluating AI-tool productivity, the difficulty of forecasting AI's effects on work and jobs, and a comparison of closed and open models on benchmarks.

## Source excerpt

Greg Wilson has noticed that lots of folks are using dodgy metrics to figure out if AI tools are worth their costs. Would you measure lines of code generated, or tickets closed? Or would you send out a survey asking whether developers feel more productive? Each of those approaches is flawed in a different way; He lists lots of common metrics, and why they are flawed. Sadly he doesn't give any suggestions on what would be better. In my view, since we cannot measure productivity, any metrics are weak evidence at the best of times. I do somewhat use one of his flawed measures: "Asking Developers If They Feel More Productive". While I acknowledge the problems he gives with this measure, I find that in an environment where decent measures are hard to find, even such a dim light is the best we have. In this situation these kinds of qualitative metrics may not be conclusive, but they are useful. ❄ ❄ ❄ ❄ ❄ Benedict Evans observes that extensive automation didn't mean the demise of professions in the past. we spent a century automating accounting: we built calculating machines, punch cards, mainframes, data processing, databases, PCs, spreadsheets, ERPs, cloud... in fact, we built half of the tech industry around automating this. Yet the number of accountants kept going up. He goes into the myriad of problems that exist when we're trying to forecast the impact of a technology on jobs. There's the much-talked-about Jevons paradox - once something becomes cheaper, people do it more, which can increase demand. Often this leads to the nature of jobs changing, even if it's called the same thing. Accountants today aren't doing exactly the same work that they did in 1970 or 1980 'but more' - they're still called 'accountants' but the job is different. New technology often starts out being used for 'the old thing but more', but it rarely ends up like that. Technologies often affect whole businesses - consider the impact of the internet on news publishing. Did anyone observing the rise