Improving Vibe Coding Efficiency and Code Quality with a Custom Skill
This post was translated from Chinese by AI. If anything reads oddly, the Chinese original is authoritative. 中文原文
Originally published at: https://www.sunai.net/t/1407 This article was written with assistance from claude opus 4.8
Quite a few people around me are frustrated with vibe coding. The models are already powerful, and tools like Claude Code are right there, but their conclusion is often the same: AI-generated code is unreliable. It runs, but they don't dare ship it. Change one thing and three others break. Reviewing it is more work than writing it themselves.
My own experience has been completely different. With the same models and tools, output quality can vary enormously. The difference usually isn't the model—it's whether you've given it your own or your team's engineering judgment. Most people treat AI as a wishing well: throw in a sentence and wait for a result. What actually gets AI to consistently produce production-ready code is turning your mental standards for "getting it right" into something it reads every time.
This article explains how I do that—the core is a code skill I wrote myself. All the examples below come from the czl-code-skill repository I maintain.
1. What a skill is, and why you should write your own
The concept of a skill is simple: a rules document with trigger conditions that AI can load on demand. Claude Code (and tools like Codex and Gemini that support the same skill-creator standard) reads it into context when appropriate, then follows its rules.
It differs from a prompt you casually send to AI in two ways: first, it's persistent—write it once and reuse it across sessions and projects, without explaining everything again; second, it has routing—rather than one long, flat document, it's an entry point plus a set of domain-specific rules. The AI reads only the parts relevant to the current task.
So why write "your own" instead of using existing guidelines online? Because programming has no universally correct standard. Naming styles, directory structures, whether to use go:embed, whether empty data should return null or an empty object, whether to use Cookie or localStorage for authentication—every team answers these differently, and many of those answers come from lessons learned the hard way.
AI doesn't know which pitfalls you've encountered. Its default behavior is to use "the statistically most common approach," and that approach is often exactly wrong for your project. Writing a skill is essentially about one thing: making my implicit engineering judgment explicit and documented, so AI follows my standards every time rather than the internet average.
Here's an example from my own setup. My skill strictly requires new projects to split the root directory by application into server/ web/ desktop/. Even with just a backend and a frontend, a flat structure is not allowed, nor is Go source code at the repository root. This isn't a universal truth; it's my preference. Without this rule, AI will almost certainly put main.go directly at the root—because that's what plenty of open-source projects do. With the rule, it organizes things neatly every time.
2. Why use a skill when you already have CLAUDE.md
This is the question I get most often. Claude Code already reads ~/.claude/CLAUDE.md (global) and CLAUDE.md at the project root (project-level). Why not just put the rules there?
That doesn't work, and the two serve entirely different purposes. Here's how I divide their responsibilities:
A project-level CLAUDE.md is a snapshot of "what this project currently looks like." It records facts, not standards: the project's architecture, directory conventions, public API paths and error codes, and counterintuitive pitfalls ("SSR is disabled for this module because…"). It's a manual for "the next AI or person taking over this project." My skill has a dedicated section defining how to write it, and the first rule is no running change logs—no "added in this update/fixed in this update," no dated lists of changes, no piles of changelog entries. I state the test plainly: if a sentence will read like an archaeological note three months from now, it doesn't belong in CLAUDE.md. I also cap the project-level file at 300 lines; if it exceeds that, it's probably full of change logs.
A skill is "my approach to writing code," which stays consistent across projects. How to decide between a major rewrite and a small fix, which edge cases to enumerate before starting, how to review my own work, when to split a long file, which naming case to use—these aren't specific to any project. They're my definition of "writing good code." They shouldn't be repeated in every project's CLAUDE.md.
There's also a very practical reason: size. If I put all my standards into ~/.claude/CLAUDE.md, it would become a monster thousands of lines long, loaded in full for every session. Besides consuming context, it would slow responses and dilute attention. A skill's routing mechanism naturally solves this—the main file is small, and detailed rules are loaded on demand. I'll cover that below.
The distinction in one sentence: CLAUDE.md answers "what this project is like"; a skill answers "how I should write code." The former travels with the project; the latter travels with me.
3. How to write a skill—a main SKILL.md plus a references directory
The structure has just two layers:
A top-level SKILL.md is the sole entry point with frontmatter. The frontmatter contains the skill's name, description, and trigger terms—the triggers matter because they determine when AI thinks to load the skill. Mine are quite specific: "write code, fix bugs, refactor, create a project, change UI, add an API, edit CLAUDE.md, Dockerfile, naming conventions, Go package structure, PWA, Service Worker, split files, file too long…" The more specific the list, the more accurately it matches.
I put only three things in the body of SKILL.md: the default architecture (the default tech stack for new projects), rule precedence (user requirements > existing project conventions > domain standards > default rules, applied in that order when conflicts arise), and a routing table from trigger terms to references. The body doesn't expand on any detailed rules; each rule gets a one-sentence summary and a link to references.
The references/ directory holds the detailed rules, with one domain per file, all plain Markdown without frontmatter. My current split includes coding-standards.md (general coding standards), go-backend.md (Go backend), nextjs-frontend.md (Next.js), frontend-design.md (UI visuals, further split into subfiles for colors/fonts/layout/loaders), pwa.md, stripe-webhook.md, and so on.
The routing table looks like this (excerpt):
| Task keywords | Required references |
|---|---|
| General standards / scope of changes / self-review / naming | coding-standards.md |
| Go backend / package structure / GORM / time zones | coding-standards.md + go-backend.md |
| Next.js / static export / Shadcn | frontend-common.md + nextjs-frontend.md |
| PWA / Service Worker / manifest | pwa.md |
When AI gets a task, it first reads the small SKILL.md, checks the routing table for the relevant references, then reads those files. A Go backend change won't load a pile of UI visual guidelines.
One strict writing constraint I've kept is this: skills and references contain only rules and reasoning, not code blocks. Backticks around field names, command names, and API names are enough; even good and bad examples are rewritten as natural-language descriptions. Code blocks go stale easily, and they encourage AI to copy specific code rather than understand the rule behind it. There are only a few exceptions for things that genuinely need to be "ready to copy"—such as the complete implementation of the PWA no-op service worker or the branded loader component. Only those cases get full code blocks.
4. How to improve skill quality—one topic per file, without long contexts
This is the most counterintuitive and valuable lesson I've learned: a skill's quality ceiling depends largely on how finely you split it, not how much you write.
Many people's instinct when writing standards is to build one comprehensive document. Thousands of lines in a single file, covering everything from naming to deployment to UI to security. The problem isn't that it's "incomplete"; it's that it can't be loaded precisely. AI either reads everything (context overload and diluted attention) or reads nothing (rules that exist only on paper).
I do the opposite: each topic or specific requirement gets its own file. Color standards, typography standards, responsive design standards, layout component standards, loader standards—I split them all out. frontend-design.md is just a main index, a global token list, and a self-check checklist. For colors, read 1-color.md; for layout components, read 4-layout-and-components.md.
This split has three benefits, in order of importance:
- On-demand loading keeps context clean. When AI changes colors, it loads only the color file instead of drowning in thousands of unrelated rules. The cleaner the context, the more accurately it follows the current rule. This is the core benefit of splitting files—you're not organizing documents; you're controlling what AI actually takes in each time.
- Rules don't conflict with each other. Related rules live in one file, avoiding "naming conventions repeated in three files, each contradicting the others." Use relative links for cross-references instead of repeating content.
- Lower maintenance costs. Changing a color rule means editing only
1-color.md, not hunting through a giant file.
I use the same standard for deciding whether to split a file as I do for code—based on responsibilities, not line count. A references file should cover the rules for just one domain. If a file covers both backend error handling and frontend animations, it should be split, even if it is only 200 lines long. Conversely, a file covering the complete Go backend conventions can be 500 lines long and still stay intact, as long as it has a single responsibility and removing any section would compromise its completeness.
Incidentally, I also put this principle of "split by responsibility, not by line count" into the skill to constrain the code AI writes. The hard rule in the skill is that each file or function must have no more than one independent responsibility. Line counts (file 400 / component 300 / function 80) are only signals to stop and evaluate; the conclusion can absolutely be "single responsibility, no split needed." But if a file exceeds 600 lines / a function exceeds 200 lines and the decision is still not to split it, the reason must be documented in a comment at the top of the file—if you cannot explain why, it probably should be split. The rules follow their own rules, and that consistency makes the whole system more credible.
5. After installing the skill, how do you ensure it actually gets invoked?
Installing the skill at ~/.claude/skills/czl-code-skill is only the first step. Installed does not mean used—AI makes its own judgments and may skip loading it because "this task is too small to need it." I learned this the hard way: the skill was installed, but when it edited a small file, it never read the skill at all. The output still followed its default style.
The solution is to explicitly require loading it in ~/.claude/CLAUDE.md, making "when the skill must be loaded" a hard rule too. My global CLAUDE.md contains this section (this is what I actually use):
Mandatory loading of czl-code-skill (hard rule)
Violating this section is a hard error, not a "style issue." This section takes priority over any self-assessment such as "the task is too small," "I already know this," or "a quick answer is enough."
Trigger condition: If the current task involves any of "code / projects / configuration / engineering documentation," you must load
czl-code-skillthrough the Skill tool and use SKILL.md to route to the relevant references before taking any further action.
There are a few key design choices here, each driven by a real problem:
- Judge by what the task actually involves, not keyword matching. I did not write "load it when you see the word 'refactor.'" I wrote "any reading/writing/modifying/deleting of source code, configuration, or documentation within a project requires loading it." AI is good at finding loopholes: if you list keywords, it recognizes only those words.
- Define an explicit list of exceptions. Pure chat, pure translation, simply looking up command usage, or operations targeting something entirely outside the project directory—only these cases may skip loading it. Fix the exceptions in writing and require loading for every other ambiguous case. I even added the sentence "always load in ambiguous cases" to block it from inventing excuses.
- Block the "task is too small" excuse. I explicitly wrote that "all rules in the skill apply equally to small changes; no task is too small to require loading it." AI's favorite shortcut is "a change this small probably does not need all that ceremony," so it must be explicitly prohibited.
- Check the routing again when switching domains. If a session moves from writing backend code to tweaking the UI, it must revisit the routing table in SKILL.md and load the references for the new context. It cannot use backend rules to modify the frontend.
- Provide a fallback if loading fails. If the Skill tool is unavailable, fall back to a concise set of rules in CLAUDE.md, and require AI to explicitly state in its response, "Could not load the skill; using fallback rules." That lets me immediately see that the output may not fully meet my standards.
This is the layer many people miss. Writing a skill is not enough. You need higher-priority global instructions to ensure it actually gets read; otherwise, however well written it is, it is just decoration.
6. A skill must include these three rules—enumerate edge cases before acting, check yourself before acting, and always review afterward
Everything above covers structure and loading mechanisms. But what really determines code quality is the rules inside the skill. If I could keep only three, I would keep the following three—they directly address the root causes of "why AI-written code is unreliable."
AI has two very consistent failure modes when writing code: first, it follows only the happy path you describe, leaving edge cases and error branches to chance; second, it delivers as soon as it finishes, never looking back at what it wrote. Together, these explain the whole problem of "it runs, but it is not reliable." These three rules specifically target those two problems.
Before acting: enumerate edge cases and branches. My skill requires that after receiving a requirement or bug report, AI explicitly build a checklist across six dimensions, mentally (or in a todo list), before writing anything. This is not about adding comments; it is about forcing it to think things through:
- Inputs: empty values, empty arrays, empty objects, extreme values (0, negative numbers, huge numbers, very long strings, emoji, newlines, quotes), invalid values, duplicate submissions, concurrent writes to the same resource
- State: not logged in, logged in but unauthorized, expired session, empty upstream response, deleted dependency resources, null fields, simultaneous operations across multiple tabs
- Flow branches: success, business failure, system failure (timeouts, 5xx, database disconnections), user cancellation/refresh/back navigation midway through, rollback after partial success, whether retries are idempotent
- Concurrency and timing: whether a double-click creates two records, multiple users competing for the same resource, implicit assumptions about the order of asynchronous callbacks
- Contract impact: whether this change affects API fields/error codes/Cookie/table schemas, and whether existing callers need corresponding changes
- Structure: first use Read to inspect the target file, decide where the new logic belongs, determine whether the file already handles several unrelated responsibilities, and whether it should be split before adding more code
The checklist does not require code to handle every item, but it does require an explicit decision for each one—handle it, explicitly ignore it, or mark it not applicable. Anything marked "ignore" or "not applicable" must be called out in the delivery notes so I can verify it. This rule directly shuts down the biggest source of bugs: "AI starts coding along the happy path."
After acting: completion review. Once a requirement is implemented, AI is not allowed to deliver immediately; it must first walk through the changes from a reviewer's perspective. I listed nine checks; here are the key ones:
- Reread the diff; do not rely on memory. Read through every modified file to catch stray debug code, leftover TODOs, commented-out dead code, missed edits in copied code, and unused imports. Also revisit the splitting signals to check whether unrelated logic has been stuffed into a single-responsibility file.
- Trace the call chain. Follow the path from the entry point (route/button/command) down to the underlying layer (database/external API), confirming that inputs and outputs at every step match expectations and that errors propagate correctly.
- Revisit the pre-work checklist. Are the edge cases and branches listed earlier actually handled in the code, rather than merely "thought about"?
- Check error paths, concurrent reentry, rollback switches, external contracts, and test validation one by one.
- Launch a subagent to review complex changes. For changes spanning multiple files or affecting core modules or external contracts, launch an independent review subagent after the self-review. Use the perspective of a "reviewer without prior context" to catch blind spots you cannot see yourself.
Any issues found during self-review must be fixed in the current round, not left as TODOs for the next one (unless they are unrelated to the current requirement or depend on unavailable external resources).
There is another rule I include in this group: deciding the scope of a change. AI has a troublesome default tendency—it treats "changing the fewest lines" as the safest option, sidestepping the real pain points and continuing to patch bad code. My skill explicitly states that "prioritize the existing project" means respecting its architecture, not defaulting to the smallest change. Every nontrivial change requires a scope assessment: first identify structural pain points (duplicated logic, incorrect abstractions, changes in one place cascading into many others), then assess "local patch vs one-time refactor," and make a clear recommendation rather than laying out both options and asking me to choose. I make the final call. When the same logic appears in three or more places, an abstraction has drifted from actual usage, a module repeatedly causes problems, or the user describes a pain point as "every time, I have to…"—these signals should favor recommending a refactor.
These three rules (enumerating edge cases, completion review, and scope decisions) are the soul of the entire skill. Architecture and naming conventions determine whether code looks good; these three determine whether it is reliable and safe to deploy. If you only have time to write three rules in your skill, write these three.
7. Keep updating the skill—write down anything worth preserving as soon as you discover it
The last point, and the easiest to overlook: a skill is not something you write once and leave alone. It is alive.
My workflow is this: whenever something "worth preserving" comes up while working with AI, I add it to the skill. What counts as worth preserving? The standard is simple—I will encounter it again in another project or session, and if I do not write it down, AI will make the same mistake or miss the same judgment next time.
A few real examples, all added to the skill later:
- Once, AI used
go:embedin a new project's Go service to embed the frontend build output into the binary. I had not asked for that; it made the decision on its own. I realized this was a default behavior that would recur, so I added a rule to the skill: Go must serve the frontend fromout/usinghttp.FileServer;go:embedis prohibited. Distribution as a single binary with zero dependencies is the only exception, and it must be documented in CLAUDE.md. - While handling Stripe webhooks, I ran into cross-project event interference—different projects shared one Stripe account, and their callbacks received each other's orders. This is a very subtle trap, so I created a dedicated
stripe-webhook.mdcovering "mandatory project identification via metadata.app, filtering by project at the entry point, signature verification and return-code semantics, and idempotency." The next time any project integrates Stripe, it will avoid the same trap. - The PWA rules also evolved gradually. At first, I simply noticed that Service Workers with caching strategies caused all kinds of caching problems on static sites. Later, I settled on "deploy only a cache-clearing no-op SW + manifest by default," and even included the complete implementation of the no-op sw.js in the skill to ensure consistency every time.
Behind almost every rule is a real mistake I made, or a real case of "the AI's default behavior wasn't what I expected." The value of a skill is this: I only make the same mistake once, and it remembers for me after that. This is the biggest difference between a skill and a one-off prompt—a prompt is a consumable; a skill is an asset that compounds.
I also have rules for maintaining it, so it doesn't turn into a sprawling log: before adding a rule, decide which references file it belongs in and extend an existing section; only create a new file if it truly doesn't fit anywhere. Don't duplicate similar rules across files; cross-reference them with relative links. Keep the SKILL.md main index concise, and put detailed rules in references. After making changes, always run the sync script to push them to each tool's skill directory; otherwise, the version actually loaded locally is still the old one.
Final Thoughts
Taken together, these seven points follow a simple logic:
AI is unreliable not because the model is dumb, but because it doesn't know your standards, and there's no mechanism forcing it to check against them. A skill does exactly those two things—make your engineering judgment explicit as rules it reads every time, then use global instructions to force it to read them every time.
Write your own skills, split them by area, and load them as needed. Build in the core rules: "exhaustively consider edge cases before starting, always review your own work afterward, and recommend refactoring when it's warranted." Then keep feeding lessons from each mistake back into them. Do this, and you'll find that AI-generated code isn't just fast to write—it's reliable enough that you actually feel confident shipping it.
vibe coding isn't about outsourcing your thinking. It's about turning the standards in your head into the AI's standards.
Comments 0