What your old code is worth
An agency's finished repositories were written off as sunk cost. Original, human-written, private code with real commit history has become valuable as training and evaluation material for AI labs. Licensing is non-exclusive, so you keep the code. It only applies to code you actually own, so the first step is reading your contracts.
Every agency that has been around for a decade has a graveyard. Internal tools nobody uses. A product that never launched. A framework the founders wrote in year three. Client projects that ended long ago, sitting in private repositories that cost a few dollars a month to keep.
On the books, all of it is worth nothing. The hours were paid for or written off years ago. I thought of mine the same way until the economics of AI gave that code a second use. You kept the IP on most of what you built. Here’s what it’s worth, and how to find out whether it is yours to license.
Why finished code became an asset
Models that write code are trained on examples of code and then tested on problems they have not seen. Both steps need material.
The public supply has been used. Open source repositories have been read by every serious model. That creates two shortages. For training, labs want code that adds something new. For evaluation, they need code the model has definitely never seen, because testing a model on problems it memorized tells you nothing.
Private repositories from working software businesses answer both. They are original. They were written by people solving real commercial problems, with the compromises and corrections that implies. They carry a history of how the code changed, commit by commit, and that history records how real engineers work. And none of it was ever public.
There is a further reason this supply is finite. A growing share of new code is produced with AI assistance. Code written entirely by people, before these tools existed, cannot be made any more. If your agency has years of it, you hold something that is no longer being produced.
As long as models need unseen, human-written code to learn from and be measured against, there is demand for what sits in your archive. How long that lasts, nobody knows.
You keep the code
The structure matters, because owners assume the worst when they hear this idea.
A licensing arrangement of this kind is non-exclusive. You grant permission for the code to be used for training and evaluation. You keep ownership. You can keep using the code, maintain it, build on it, and license it again elsewhere. Nothing is handed over in the sense of losing it.
A building you rent out is still your building. Check that any agreement you are shown says “non-exclusive” in plain words and does not include an assignment of rights. If it does include one, it is a different deal and you should treat it as such.
It only applies to code you own
This is where most of the work is, and where agencies get it wrong in both directions. Some assume they own everything they wrote. Others assume they own nothing. The truth is in the contracts, project by project.
Sort your repositories into four groups.
| Group | Typical ownership | Can you license it? |
|---|---|---|
| Internal tools, abandoned products, experiments | The agency, provided staff and contractor agreements assigned their work to you | Usually yes |
| Reusable libraries and frameworks you brought to client projects | The agency, if your contracts carved out pre-existing and background IP | Usually yes, check the carve-out |
| Client work under contracts where you retained ownership and licensed the client | The agency, subject to confidentiality terms | Possibly, read the confidentiality clause |
| Client work under contracts that assigned everything to the client | The client | No |
Three checks decide which group a repository belongs in.
The client contract. Look for the words “assigns”, “work made for hire” and “all right, title and interest”. If the deliverable was assigned to the client on payment, it is theirs. Many agency contracts assign the bespoke deliverable and keep the agency’s pre-existing tools and general components. That carve-out is worth a great deal now. The full explanation is in who owns the code, and the wording itself is covered in the IP clause in a software development agreement.
The confidentiality clause. Even where you own the code, a confidentiality obligation may restrict what you can do with anything that reveals the client’s business. Ownership and confidentiality are separate questions. Answer both.
Your own people. Employees’ work generally belongs to the employer in the US, with variations elsewhere. Contractors are different: without a written assignment, a contractor may own what they wrote for you. This catches agencies that used freelancers on a handshake. In the UK and EU the default rules for contractors point the same way, and moral rights can add a wrinkle. If the paperwork is missing, get it signed before you do anything else.
When ownership is unclear, treat the repository as not yours until someone qualified says otherwise. Licensing code that belongs to a client is a breach of contract and a quick way to lose a reputation that took fifteen years to build.
What makes a repository more or less valuable
No figures here. Value depends on the specific repository, and anyone quoting a number before looking at one is guessing. The qualities that move it are consistent, though.
Originality. Code your team wrote counts. Forks of open source projects, vendored dependencies, copied snippets and generated boilerplate do not. A repository of 200,000 lines with 180,000 lines of third-party packages is a 20,000-line repository.
Human authorship. Code written by people, particularly before AI assistants were in common use, is the scarce kind. A repository’s dates matter for that reason.
Depth of history. A long, genuine commit history with many commits by several authors over months or years shows how the software evolved: the bug, the fix, the refactor, the review. A single “initial commit” containing the finished code carries far less information.
Language and domain scarcity. Common web stacks are abundant. Less common languages, older enterprise stacks, embedded code, scientific computing and specialist industry software are thinly represented in public data, and tend to be more sought after for it.
Substance. Real business logic, tests, and evidence that the software ran in production count for more than tutorials, prototypes and scaffolding.
Clean ownership. A repository with a clear chain of title, with contracts and assignments on file, is usable. One with a question mark over it is worth very little to a careful licensee, whatever its quality.
No secrets, no personal data, no client data. API keys, credentials, customer records and database dumps in a repository, or anywhere in its history, are a problem that must be dealt with before anything else happens.
If you want a rough sense of where a repository stands before talking to anyone, there is an independent tool for it: a repo value calculator that estimates from characteristics like size, age, language and history.
Do not tidy the history
The instinct of every engineer, on hearing that someone will look at an old repository, is to clean it up. Squash the messy commits. Rebase. Rewrite the embarrassing messages. Start a fresh repository with the final code.
Do none of that. The history is a large part of what is being assessed. The messy sequence of real work, with its false starts and fixes, is the information. A squashed repository has thrown away the thing that made it distinctive.
Also avoid running the old code through an AI tool to modernize or reformat it first. That turns human-written code into something else.
Secrets and personal data do have to be addressed, and that sometimes involves the history. Do it deliberately and narrowly, and take advice before you start, since a careless purge wrecks the commit record you are trying to preserve. Rotate any credential that was ever committed, whatever else you do. That is basic hygiene whether you license anything or not.
How to take stock
A practical afternoon’s work for an owner or a technical lead:
- List every repository the business controls, across every hosting account, including archived ones and the old server in the cupboard.
- Record the basics for each: what it was, when it was written, main language, rough size, number of commits and authors.
- Assign each to one of the four ownership groups in the table above. Pull the contract for anything that was client work.
- Flag the risks: secrets, personal data, client-confidential material, contractor code without an assignment.
- Rank what is left by the value factors: originality, history, scarcity.
- Leave the repositories exactly as they are until someone has assessed them.
Most agencies that do this find a smaller pile than they hoped and a more valuable one than they expected. The internal tools and retained libraries are usually the core of it.
Once you know which repositories are yours, the next step is an assessment of what they are worth under a non-exclusive license. See what your repositories are worth.
Fix the contracts going forward
Whatever you find in the archive, the lasting lesson is about the contracts you sign from now on. An agency that assigns everything to every client by default ends each project owning nothing. An agency that assigns the bespoke deliverable and keeps its background IP, tools and reusable components builds an asset with every engagement.
Clients rarely object to a sensible carve-out, because they get what they are paying for: full rights to use and modify their system. If a client is nervous about being locked out, source code escrow or a broad license usually settles it. Put the clause in your standard agreement and stop negotiating it from scratch. The wording is in the IP clause in a software development agreement.
Where this fits
Code licensing is a one-time or occasional return on work already done. It will not replace a services business. It is a useful line in a year when build revenue is under pressure, and it costs little beyond the time to check ownership.
It fits a larger shift. Cheap code generation is pushing down the price of routine implementation, which I cover in is software engineering dead. The same shift has raised the value of code that was written the slow way. Agencies sit on both sides of that change. The work of building new revenue lines to match is in the AI consulting business and the future of software engineering, and everything else on the subject is in the AI hub.
Open the archive and read the contracts. The code you stopped thinking about years ago may be the one asset in the business that gained value while you were busy with client work.
Common questions
- Can an agency license code it wrote for clients?
- Only code the agency owns. If the contract assigned all rights in the deliverable to the client, that code belongs to the client. Internal tools, abandoned products, reusable libraries you retained, and work under contracts where you kept ownership are the usual candidates. Read each contract before assuming anything.
- Do I lose my code if I license it?
- No, provided the license is non-exclusive, which is the normal structure for this kind of arrangement. You keep ownership and can go on using, modifying and licensing the code. Confirm the non-exclusive wording in any agreement before you sign.
- Why would old code be valuable to AI labs?
- Models that write code learn from examples and are tested against examples. Public code has been used heavily already. Private code written by people, with a history showing how it changed over time, is material the models have not seen, which makes it useful for both training and evaluation.
- What makes a repository more valuable?
- Originality, a deep and genuine commit history, clean and provable ownership, and languages or domains that are thinly represented in public code. Forks, generated code, vendored dependencies, squashed history and anything containing secrets or client data reduce the value or rule the repository out.
- Should I clean up a repository before having it assessed?
- Do not squash, rebase or rewrite the commit history. The history is a large part of what is being valued. Dealing with secrets and personal data is necessary, and it should be done carefully and with advice, since crude cleanup can destroy the history you are trying to preserve.