g3ddy
AI Strategy8 min · 9 Oct 2026

Software engineering now: what's changing and what to master

Vibe coding proves an idea, but production needs specs, test harnesses, and ownership. A year of signal on what agentic engineering actually looks like.

Geddy
Geddy
Senior Web Engineer / Lead
0:00
A photorealistic modern office desk scene shot from a low wide angle, a single monitor displaying layered dashboards of passing test suites and security audit checklists, soft blue monitor glow illuminating a focused developer's hands on keyboard, blurred background bokeh of a dim office at night, shallow depth of field, cinematic moody lighting, wide 16:9 framing

AI-generated content. This article was generated by AI from material I collected and has not been fully reviewed by me. It may contain errors — verify anything you rely on. More in the disclaimer.

TL;DR30 sec
  • Vibe coding gets you a demo; production still demands error handling, rate limiting, real backends, and someone who can defend every line
  • The job shifted from writing code to designing the harness — specs, test gates, review posture, parallel-agent workflows — that lets agents ship code you're accountable for
  • Verification beats generation: write the failing test first, gate agent output behind CI, and run a structured security audit inside the loop instead of after it
  • Pick one review posture (live-approve or end-of-run PR) and stick to it — randomly mixing the two is how code quality quietly drifts
  • Score models on your own repo before trusting them; a cheap model matched a premium one on issue coverage at 3% of the cost
  • "The AI did it" ends careers in code review, and the hiring bar has moved from discussing agentic systems to having shipped one under real load

Stop prompting better. Build the system that makes the agent's output verifiable.

The role isn't "write code faster with AI" anymore. It's designing the system that lets AI agents write, test, and ship code you're still accountable for.

I've been saving screenshots for a year — threads, repos, job postings, benchmarks, arguments. This is the engineering-craft signal pulled out of that pile. Claude Code setup itself is covered everywhere else. This is about the craft around it.

TL;DR

  • Vibe coding proves an idea. Production still needs error handling, rate limiting, and real backends.
  • "Agentic Engineering" (Karpathy's term) is now a named discipline: spec-first, fixed review posture, test harness, parallel-agent design.
  • Spec-first tooling went mainstream — GitHub's Spec Kit hit 126k stars and works across 25+ agents.
  • Verification, not code volume, is the differentiator. TDD-first skill libraries have crossed 244.7k stars.
  • Harness and loop engineering — the system around the model — is where AI pilots actually die in production.
  • Security review is moving into the agent loop, not staying a post-hoc gate.
  • "The AI did it" is no longer an acceptable answer in code review. Ownership is the baseline.
  • The hiring bar moved from "can talk about RAG/agents" to "can ship and debug something under real load."

What's changing

Author → orchestrator. Engineers are increasingly managing dozens of parallel agents rather than writing every line themselves.

Ad-hoc prompting → fixed review posture. Teams that randomly mix real-time approval with end-of-run PR review produce inconsistent code. The fix isn't picking the "right" one. It's picking one and sticking to it.

README → constitution file. CLAUDE.md-style rules files are becoming executable architecture contracts the agent reads on every run.

Security as a gate → security in the loop. Pattern-catching, diff review, and commit-time validation now run inside the coding session itself.

Tooling speed as an assumption → a measured cost line. One team quantified 1,140 engineer-hours a year lost to a slow linter before migrating. That's a budget item, not a preference.

"Can discuss agentic systems" → "has shipped one." Interviewers are filtering for production experience, not vocabulary.

Design vs build → design engineers who ship. The line between designing an interface and building it is dissolving.

New tendencies and direction

Spec-first / agentic engineering. Write the spec and architecture rules before the agent touches code. Four pillars: spec-first, fixed review posture, test-harness-first, parallel-agent design.

Verification over writing code. TDD-first skill sets — failing test, fix, next feature — and dedicated review/audit skills are replacing "just generate the feature."

Harness and loop engineering. Prompt engineering is the message. Context engineering is what the model sees. Harness engineering is the system that gathers, acts, and verifies. Most failed AI pilots never built the harness.

Loops as infrastructure. A loop — automations, worktrees, skills, connectors, sub-agents, plus memory of what's done and what's next — is the floor above harness engineering.

Trust ledgers and budgets for autonomy. Production-grade agent systems earn unattended-run permission only after logged pass-rate evidence, with hard spend caps enforced in code, not policy.

Cost-aware model selection. Pick models by scored task and budget, not reputation. In one audit, a cheap model matched a premium one on issue coverage at 3% of the cost.

Instruction architecture as a transferable skill. The public system prompts behind Claude, Cursor, and others share the same shape — clear roles, explicit "do not" lists, a stated working model. Worth reading directly.

Practices that matter now

Practice Why now How to apply
Spec/constitution before code Agents don't know your conventions or business logic Write a CLAUDE.md / spec-kit constitution before delegating any task
Fix your review posture Randomly mixing live-approve and PR-review modes is how quality drifts Choose real-time or end-of-run review per task type and stay consistent
TDD-first with agents Agents guess at unclear requirements without a failing test to anchor them Write the failing test, let the agent make it pass, feature by feature
Structured security audit before merge A 3-layer review cut security PR comments 30–40% internally at Anthropic Run a 5-area audit (secrets, deps, auth, OWASP, infra) before every merge
Own every AI-assisted PR "I don't know, Claude did it" is not a defensible answer Be able to explain and defend every line you submit
Score models before trusting them Reputation isn't performance; cost and coverage vary wildly Benchmark 2+ models against known issues in your own repo
Measure your tooling cost Small per-commit delays compound into real engineer-hours Profile linter/CI steps; migrate when the ROI is proven, not assumed

Architecture and frontend insights

Architecture patterns scale with seniority. Junior needs about 4 — layered, client-server, monolith, CRUD. Middle needs 12 — CQRS, event-driven, API gateway, outbox. Senior needs 20 — saga, strangler fig, sharding, service discovery, and crucially, knowing when not to use a pattern.

Five API paradigms now coexist on a single stack. REST, GraphQL, and gRPC connect software to software. MCP connects software to a model. A2A connects agents to agents.

Generative UI is emerging as "frontend for agents," with three patterns: pre-built components the agent picks from, a declarative schema the app maps to components, or open-ended HTML rendered in a sandbox. A streaming alternative to JSON-based generative UI (OpenUI Lang) claims 67% fewer tokens and 3x faster rendering, with Zod-typed contracts and zero arbitrary code execution.

Anti-"AI slop" design tooling is a new category — rule sets that ban the generic purple-gradient, centered-hero look, plus agent-ready design systems like Facebook's Astryx.

And the unglamorous one: ESLint + Prettier → Biome cut linting from 55.6s to 0.89s per commit. 98% faster, directly reproducible, measurable.

Proof points matter. On my own site: Lighthouse Performance 96–99, Best Practices 100, SEO 100, CLS as low as 0.001–0.036. That's a screenshot anyone can verify in thirty seconds.

Competencies to build

Competency What "good" looks like How to prove it
Agentic engineering Specs before delegating, fixed review posture, output gated behind a test harness, tasks designed for parallel agents Publish a repo with a constitution file + CI harness gating agent-written code
Architecture judgment Knows when not to use a pattern — microservices too early, CQRS without real need Write up a trade-off decision, including the alternatives you rejected
Verification-first workflow Failing test exists before the implementation does Show a PR history where the test commit lands before the feature commit
Structured security review Consistent multi-area audit as part of the loop, not after Attach a real audit output to a public PR
Cost/model literacy Chooses models by scored task and budget, not brand Publish a benchmark comparing cost and accuracy across 2+ models on your repo
Ownership under review Can explain and defend every line in an AI-assisted PR Never answer "the AI did it" — narrate the actual decision

Traits of engineers who win now

They ship systems that survive real load — error handling, rate limiting, real backends — not just working demos.

They orchestrate multiple agents in parallel instead of pairing with a single assistant all day.

They build the harness once — specs, tests, review gates — instead of re-prompting from scratch every session.

They take full, explicit ownership of AI-generated code in review.

They treat security and cost as engineering constraints they design for, not someone else's job.

Tools worth adopting

Tool What Use it for
GitHub Spec Kit Spec-first CLI: /constitution → /specify → /plan → /tasks → /implement. 126k★, 25+ agents Forcing a plan before an agent touches code
"Skills For Real Engineers" 25+ Claude Code skills for requirements, specs, TDD, review, architecture. 244.7k★, built on Pragmatic Programmer / DDD principles Making agents work like a senior engineer by default
Matt Pocock's Claude Code skills /grill-me, /write-a-prd, /tdd, /triage-issue, /git-guardrails. 46k★ Structured discipline for real work, not demos
Anthropic security-guidance plugin 3-layer review: instant pattern catch (zero inference cost), full diff review, commit-time validation Shifting security review left, inside the session
Biome Rust linter/formatter, drop-in for ESLint + Prettier Cutting CI/commit-time linting cost at team scale
direnv Per-directory env var loading, cascades from parent .envrc Managing API keys across multiple Claude Code/MCP projects
Obscura Rust headless browser: ~30MB RAM, ~85ms page loads, per-session fingerprint randomization, Puppeteer/Playwright drop-in Scraping and browser automation for agents
LlamaIndex LiteParse v2 Rust-rewritten PDF parser; 457 pages / 100MB in 0.777s Fast local document parsing in AI pipelines
taste-skill / Hallmark / Astryx Anti-slop design rule sets and an agent-ready open-source design system Avoiding the generic AI-generated UI look
LLM Anonymization proxy Transparent proxy between Claude Code and the API; strips secrets/hostnames before they leave your machine Claude Code on sensitive client work

Selling your engineering + AI skills

Publish before/after metrics you already have. A Lighthouse audit is instant, verifiable proof of craft — and it costs you nothing to produce.

Turn a real project into a case study. A public relicensing decision — say, moving a repo to PolyForm with a personal-use carve-out — shows architecture and legal literacy in one artefact.

Share your actual Claude Code session: the todo list, the diff size, the PR. Process evidence reads stronger than a finished screenshot alone.

Geddy
Geddy
Senior Web Engineer / Lead

Engineering leadership • AI innovation • Product thinking. 20+ years of web engineering, from independent contractor to engineering leader. Passionate about developer experience and product engineering.

Follow on LinkedIn
Keep reading