The split between 'serious work' AI and 'just vibing' AI happened when stakes started sorting the tools, often without formalizing it. This behavior emerged as a trust hierarchy, where tasks with higher risk are assigned to more reliable models. The division is evident in task routing, not branding, and reflects a new work norm.
Verdict: The split happened when tasks were sorted by stakes.
The split happened when stakes started sorting the tools
Claude Opus for planning and UI bugs, Gemini 3.1 Pro for routine work. That is the split in plain view, and according to dev.to it showed up as behavior before it became a stated rule. People now keep a “serious work” AI and a “just vibing” AI, and AI means machines emulating human intelligence. The key change is that we started routing tasks by risk and consequence, often without ever announcing that we were doing it.
| Task type | Model people give it |
| Planning, UI bugs without an obvious cause | Claude Opus |
| Routine work, smaller changes | Gemini 3.1 Pro |
This is a trust hierarchy disguised as convenience.
The strongest evidence is task routing inside one workflow. Antigravity is used as the AI IDE, which means the model choice happens in the middle of real work rather than as a thought experiment. Hard, ambiguous tasks go to the model treated as safer; routine edits go to the model treated as good enough.
If you run this at work, the practical consequence is simple: your team already has two standards for acceptable error, even if nobody has written them down. Watch where people send planning, debugging, and low-stakes edits, because that is where the split became real.
The evidence is in the task routing, not the branding
Claude Opus for planning and UI bugs, Gemini 3.1 Pro for routine work, then Claude again for names and urgent document edits is the cleanest evidence that people now keep a serious-work AI and a vibing AI, according to dev.to. That split came from repeated risk choices inside one workflow, not from a formal policy.
- Planning and UI bugs with no obvious cause go to Claude Opus.
- Routine work and smaller changes go to Gemini 3.1 Pro.
- Names, ideas, and urgent doc edits that need to sound right fast go to Claude.
Trust is being allocated by stakes, task by task.
The same person is accepting different error budgets in different lanes.
That is why this is a behavioural hierarchy. Gemini 3.1 Pro hallucinates more than Claude Opus in that context, yet it still gets the smaller changes because the cost of being wrong is lower there. Gemini also gets calculation-heavy problems because it has never given a wrong answer there.
If you run this at work, watch your own routing before you read any vendor claim. The model you reach for when a bug is weird or a document is urgent is the one you trust under consequence, and the one you tolerate for routine churn is your good-enough lane.
The best objection is that this is just normal tool choice
Antigravity is used as an AI IDE, according to dev.to, and that matters because interfaces often decide behavior before culture does.

Claim: The strongest objection is simple: this looks like ordinary tool choice, the same way people use one app for spreadsheets and another for chat.
Counterargument: People have always split work across specialized software, and model switching can follow subscriptions, usage caps, or product packaging rather than any new social meaning. If one system sits inside your editor and another lives in a friendlier chat window, people will route tasks by convenience alone; dev.to even gives a plain example in Antigravity as the work surface.
That is a serious objection.
Rebuttal: It still falls short because the pattern people describe is not only “this tool is here” but “this tool is for higher-stakes work,” which is a ranking, not mere placement. Dev.to reported, “Somehow, I just ended up with a ranking in my head,” and if you run AI in production, that is a trust policy whether you wrote it down or not.
That objection misses how quickly the hierarchy hardened
Early 2020s is when generative AI came to the forefront: systems that create text, images, audio, and code that often pass for human-made, and that is when the split hardened into a work habit. People did not merely get more options. They began assigning different systems to different levels of risk and seriousness, often without deciding to, because once a task can cost time, money, or credibility, acceptable error drops and trust gets sorted fast.
| Wider context | Figure |
| Early 2020s | Generative AI came to the forefront |
| Early-career employment in AI-exposed roles | 13% relative decline |
| US jobs projected lost by 2030 | 6.1%, about 10.4 million roles |
Personal routing turns into policy the moment a manager asks which model is safe enough for client work.
According to stonybrook.edu, early research found a 13% relative decline in employment for early-career workers in AI-exposed roles. The same source reports a projected 6.1% loss of US jobs by 2030 from AI and automation, about 10.4 million roles.
That is why this taxonomy matters. Under pressure, teams do not sort tools by vibes alone. They sort by where errors are tolerated, where review is mandatory, and where blame lands. If you run this in production, watch the approval path: the “serious” model is the one people trust when mistakes become expensive.
What this means for anyone choosing an AI at work
13% is enough to make this practical, not philosophical: according to Stony Brook University, early research found a 13% relative decline in employment for early-career workers in AI-exposed roles. Once work is under that kind of pressure, teams stop talking about “favorite models” and start sorting them by where error is tolerated.
**Why it matters:** If you already switch models by task, you have a trust policy whether you wrote one down or not. The useful move is to name it, because unnamed policies fail exactly when stakes rise.
That means choosing an AI at work gets simpler and stricter.
Put the split where the risk already is
A task belongs in the serious work lane when:
- an error changes a decision, a number, or a customer outcome
- the output will be reused by someone else without rechecking
- you know from experience that one model is less reliable for this kind of work, as with Gemini 3.1 Pro hallucinating more than Claude Opus in one comparison
The opposite lane still has value.
I use that lane for speed, drafts, and low-cost exploration, because “good enough” is a real standard when the cost of being wrong is low. But for calculation-heavy work, one practitioner reported Gemini had never given a wrong answer in their experience.
That is now a work norm, not a hobbyist quirk.
The split is real, and work made it stick
Q: When did the split between a serious-work AI and a vibing AI actually happen? A: It happened when teams started routing tasks by stakes inside real workflows. The early 2020s gave people enough capable models, and work pressure turned preference into a trust hierarchy fast.
Q: Is this just ordinary tool specialization? A: No. Ordinary specialization explains placement, but this pattern is a ranking by acceptable error. The model used for weird bugs, planning, and urgent docs is the one trusted under consequence.
Q: Why does this matter at work? A: Because an unwritten routing habit is already a policy. If you run AI in production, mistakes, review load, and blame will follow the lane you assign to each model.
Q: What should a team do next? A: Write down the routing rules you already use. Mark which tasks require the serious lane, which can use the good-enough lane, and review that map before a vendor pitch rewrites it for you.
Sources
- 1. [We All Have a "Serious Work" AI and a "Just Vibing" AI. When Did That Happen? - DEV Community]( Serious Work
