themisf.it

What good is a skill if an AI is going to do it eventually?

Earlier this year, I finished a bachelor's in IT. I have been working in the field for a long time, and the program taught me nothing I did not already know. That sounds like a knock, but it really isn't. It would be more troubling if someone in my position had found it revelatory, given the years behind me and a first technical certification passed when I was barely a teenager. Competency-based education assumes you are demonstrating what you have rather than acquiring what you lack, and the assumption held. I spent a couple of weeks proving I could do things I have been doing professionally for two decades.

Set that next to the dominant conversation in my field, which is about how many of those things are about to be performed by software, and you have a small puzzle. A rational person does not invest in certifying a capability whose market value is collapsing. Either I did something irrational, or the thing I bought was not what it appeared to be.

The same puzzle shows up constantly right now in a broader form. Should I learn to code if the models write code. Should I pursue the certification if the work is being automated. Should I develop the skill if the skill has a shelf life measured in years rather than careers. The question arrives from people standing at the entrance to a learning investment, trying to decide whether to walk through, and the answers available to them are mostly bad.


The explanations that do not work

The first says that artificial intelligence will never replicate genuine human creativity, insight, or judgment, and that the irreducibly human portion of the work will remain. This may turn out to be true. It is also a prediction about the future stated with far more confidence than the evidence supports, and it has the structure of a claim that keeps retreating. Every time a capability falls, the boundary of the irreducibly human moves inward, and the argument survives by redefinition rather than by being right. An explanation that cannot fail is not doing explanatory work.

The second says to learn to work alongside the tools rather than compete with them. This is true and nearly useless. It offers no guidance about which skills to acquire, in what order, to what depth, or why any particular investment would pay. It describes a posture rather than a mechanism.

The third says to concentrate on the soft skills that machines cannot touch. This underestimates what the machines are already doing to persuasion, summarization, synthesis, and emotional register, and it condescends to everyone whose professional value runs through technical depth. It also gets the causation backward, as I will argue below: the interpersonal capacities that survive do so because they sit on top of domain judgment, not because they are separate from it.

All three share a defect. They are attempts to identify which categories of work are safe. The category question is the wrong question, because the thing being automated does not respect category boundaries. Something more structural is happening, and it operates underneath the level of job titles and skill taxonomies.


What a skill has always been

A skill has always purchased two different things at once. The first is the capacity to produce the output. The second is the capacity to evaluate the output. Historically these arrived together, bundled so tightly that no one had cause to distinguish them, because the only route to evaluation capacity ran directly through execution capacity. You learned to judge a database schema by building enough of them badly. You learned to recognize a bad security architecture by having your own architecture fail. You learned which vendor claims were plausible by believing implausible ones and paying for it.

Because the bundle was inseparable, we gave it a single name. We called it skill, or expertise, or experience, and we measured it in years, because years were a reasonable proxy for how many cycles of production-and-consequence a person had accumulated. Automation unbundles the two.

Generative systems supply execution capacity at close to zero marginal cost. They do not supply evaluation capacity, and the reason is structural rather than temporary. The fluency of the output and the correctness of the output are produced by different mechanisms, and only one of them is guaranteed by the process that generates the text. A model producing a plausible-sounding architecture and a model producing a correct one are doing the same thing internally. The difference lives in the world, not in the generation.

So the person holding the output needs some way to tell which one they received. That way is domain judgment, and there is no substitute for it, and it can still only be built the old way.


Why good people invest in the wrong capacity

The interesting consequence is not that evaluation capacity matters. It is that the people most likely to under-invest in it are the ones behaving most rationally.

Consider how a career allocates its scarce resources. A professional has finite time and attention for learning, and they direct it toward whatever their current role measures and rewards. Almost universally, roles measure execution. Tickets closed, features shipped, courses developed, reports produced, audits completed. These are legible, countable, and reviewable. Evaluation capacity, by contrast, shows up in the absence of problems, in the projects that were killed before they consumed a budget, in the vendor contract that was not signed. It is nearly invisible to any measurement system an organization actually operates.

A professional optimizing correctly against the incentives in front of them will therefore keep investing in throughput. They will get faster. They will handle more volume. They will be rewarded for it, accurately, because throughput is what their role was constructed to produce. And they will be pointing their entire learning budget at exactly the capacity whose market value is collapsing, for locally correct reasons, right up until the collapse reaches them.

This is not a story about complacency or about people failing to see what was coming. It is a story about a resource allocation process producing a systematically bad outcome through a sequence of individually sound decisions. Organizations do the same thing at a larger scale, which is why the wave of AI investment currently underway is so heavily concentrated in making existing processes faster rather than in building the institutional capacity to tell whether the faster output is any good.


Where the theory stops holding

A theory that explains everything explains nothing, so here are the conditions under which this one applies.

Evaluation capacity commands a premium only where the cost of a wrong answer is high relative to the cost of producing an answer. Where output is cheap and errors are cheap, nobody will pay for judgment, and they should not. Draft marketing copy, first-pass meeting summaries, throwaway scripts, exploratory analysis that will be checked by something downstream: in these domains the automation story is straightforwardly the whole story, execution value is the only value, and the correct move is to use the tools aggressively and stop thinking about it.

The premium appears where errors are expensive, delayed, or hard to detect. Financial models that inform capital allocation. Security architectures whose failure mode is a breach eighteen months later. Clinical and legal work with regulatory exposure. Infrastructure decisions that take five years to unwind. Anything where a plausible wrong answer will be acted upon before anyone discovers it was wrong.

The distinction does not track technical versus non-technical, or creative versus routine. What matters is the cost and latency of error. Locate your own work against that boundary honestly, rather than assuming your domain sits on the comfortable side of it.


The observable signature

The unbundling has a visible tell, and once you have seen it a few times you cannot stop seeing it.

Someone with framework vocabulary and no operational depth produces something polished. It reads well. The terminology is correct and current. The structure is confident. Then a person with actual depth asks one specific question about how a claim would be measured, or what the current-state baseline is, or which system the data would come from, or what happens when the assumption in the third paragraph turns out to be false. The structure collapses, because the vocabulary was never attached to anything underneath.

The diagnostic is what happens next. A person with evaluation capacity, asked a question they cannot answer, will say they do not know and describe how they would find out. A person without it will produce more vocabulary. That is the signature. Under pressure, substance produces specificity and its absence produces fluency, and the two are easy to distinguish if you know the domain and nearly impossible if you do not.

This also explains why the interpersonal skills that survive are the ones grounded in domain judgment. The ability to say "I do not know, here is how I would find out" in a room full of executives is not a communication skill in isolation. It requires knowing the shape of the question well enough to know you cannot answer it, which is a form of domain knowledge, and having enough standing to say so, which is usually earned by having been right before.


What this means if you are looking for work

Credentials still function, and the reason has almost nothing to do with capability. My degree does not make me better at my job. It makes me legible to systems that filter candidates before a human reads anything: an applicant tracking system with a hard requirement, a recruiter with three minutes and a checklist, a committee that would otherwise spend time resolving a question mark. That filtering function is orthogonal to whether a machine can perform the certified task, and it gets stronger rather than weaker as the applicant pool grows. So buy credentials for access and be honest with yourself that access is what you are buying. A certification that opens a filter is worth pursuing even if it teaches you nothing. One that teaches you nothing and opens nothing is worth skipping.

Position on judgment rather than on tools. A resume listing technologies competes against every other resume listing technologies, in a market where tool familiarity is being commoditized in real time. A resume demonstrating that you have made consequential calls, been wrong, corrected, and can articulate the correction competes in a much thinner field.

Expect the interview to be a depth probe, and understand that the probe is the entire evaluation. The specific question matters less than whether your answer has floor under it. People with real depth can go three levels down on any claim they make and will tell you where their knowledge ends. Interviewers who are any good are testing for exactly that, whether or not they would describe it that way.

And notice that the capacity to distinguish substance from fluency is itself becoming a hiring criterion. Organizations are currently absorbing an enormous volume of confident, well-formatted, professionally packaged output, some of it excellent and some of it hollow. The people who can reliably sort one from the other are scarce, and the scarcity is going to get worse before it gets better.


The larger problem

Everything above assumes evaluation capacity can be built. That assumption is currently under attack.

Entry-level work is being eliminated across a range of fields right now. The junior analyst who used to build the model. The associate who used to do first-pass document review. The new developer who used to write the boilerplate and fix the small bugs. The coordinator who used to assemble the report. Organizations are looking at that work, correctly observing that a model can produce it faster and cheaper, and removing the position.

Every one of those roles was doing something other than producing output. It was the mechanism by which a person accumulated cycles of production and consequence, which is the only known way to build the judgment that the same organization will need from that person in fifteen years. The execution work was the tuition. It looked like the product, so it got priced like the product, and now it is being eliminated on the basis of that price.

The consequence is a slow-motion supply problem that will not surface for a decade and will be nearly impossible to fix once it does. An organization that stops making seniors has to buy them, in a market where every other organization made the same decision and the supply is fixed at whatever was already in the pipeline. The price of scarce judgment in that market will be unpleasant. Worse, a meaningful portion of the judgment an organization actually depends on is specific to its own systems, its own vendors, its own failure history, and cannot be purchased at any price because it does not exist anywhere else.

Note that this is the same failure as the individual one, running at a larger scale. Developing a junior costs senior hours, which are the scarcest resource in the building. The payback arrives in five to ten years, which is past the horizon of every measurement system the organization operates. So the locally correct decision is to eliminate the role and redirect the senior hours to current throughput. Nobody is being short-sighted in a way they would recognize as short-sighted. The resource allocation process is producing the outcome on its own, exactly as it did in the individual case, and the people making the calls are responding accurately to the incentives in front of them.

What replaces it

The answer, if an organization wants one, is closer to apprenticeship than to any of the current alternatives.

This is not mentorship in the sense the word usually carries, meaning a standing coffee meeting and career advice. Apprenticeship means a junior doing real work with real consequences, under a senior who reviews the output closely enough to catch what is wrong with it and explains why, repeatedly, for years.

What makes this newly viable is the same automation that created the problem. The historical constraint on apprenticeship was that the apprentice had to spend most of their time producing volume, because the volume was needed and producing it was slow. Judgment accumulated as a byproduct, inefficiently, over a long period. Automation removes the volume constraint. A junior with good tooling can produce in an hour what used to take a week, which means the overwhelming majority of their time can now be spent on the part that was always the actual education: defending the output to someone who knows where the bodies are buried.

Concretely, that looks like a few things. The junior produces with whatever tooling is available and is then required to justify every consequential choice in the output to a senior who probes it. The junior sits in the rooms where consequences land, which means incident reviews, post-mortems, vendor negotiations, and budget defenses, not just the rooms where work gets assigned. The junior is given decisions small enough to be survivable and real enough to hurt when they go wrong, and then is walked through what went wrong afterward. And the senior's hours spent on this get counted as an output of their role rather than as overhead subtracted from it, because otherwise the measurement system will starve the program without anyone deciding to.

The boundary condition from earlier applies here too. This is expensive, and it makes sense only where the cost and latency of error are high enough to justify it. An organization whose errors are cheap and immediately visible does not need to grow judgment internally and should not spend senior time doing it. An organization whose errors are expensive, delayed, and hard to detect is currently dismantling its own succession pipeline and calling it an efficiency gain.

There is a resilience argument here as well, separate from the succession one. An organization where evaluation capacity is concentrated in a handful of long-tenured people has a single point of failure attached to each of them, and every one of those people will eventually leave. The apprenticeship model distributes that capacity across more of the organization, which is the actual meaning of institutional resilience, as distinct from the documentation exercise that usually gets the name.


So: what good is a skill if an AI is going to do it eventually?

The question assumes a skill is one thing. It is two, and only one of them is being automated. The half that produces the work is being commoditized quickly and thoroughly. The half that knows whether the work is any good is becoming the scarcer input, and it can still only be built by producing work and living with the consequences, which is the same way it has always been built.

That is the case for spending two weeks certifying capabilities you already had, and for learning things a machine can already do. You are not buying the ability to produce the output. You are buying the ability to know what you are looking at.

The organizational version of the question is what good is a junior if a model can do their work, and it has the same answer. The junior was never the output. The junior is the only mechanism anyone has found for producing the person who will be able to tell, fifteen years from now, whether the output is any good.

— Chris