The model is not the asset. The labeled file is.
The model is not the asset. The labeled file is.
You are building tools, training on your own data, and letting models write first drafts of decisions people used to make. That creates value. The law has not caught up. Waiting for it to settle is how you lose the file in a vendor clause.
People talk about outputs. The money is usually somewhere else. The labeled set you cleaned and validated is the expensive part. A competitor with your QC images or scored customer notes can copy the model faster than they can rebuild the pile. Treat that file like a secret: origin, access, and a sentence in every vendor contract that says you still own it.
What you can actually protect
Most teams fine-tune someone else's model. The base belongs to OpenAI, Anthropic, or Google. Their terms decide what you own of the tune. Read that before you start, not after. Patent the process if it is novel: the steps, not the weights. USPTO will not name a model as inventor. Copyright will not register a work that is only a model. Human prompts, selection, and review are what make a claim.
Write the human part down as you go. After-the-fact notes are weaker. A provisional filing buys a date while you finish the real application. The queue is getting crowded.
Vendors, trade secrets, this quarter
Colab, Hugging Face, Claude, ChatGPT, Make, Zapier: each has a different story about your inputs. Some train on them. Paid plans sometimes do not. Inventory the tools, the data that went in, and the clause. That list usually finds the leak.
Trade secret works if it has value and you actually keep it quiet. One person pasting the training set into a public chat ends it. Access control, NDAs that mention AI assets, and a rule people follow. Audit the stack. Classify the data. Teach inventors to write down their part. Renegotiate the vendor line that hands them the IP. Do it this quarter, not when the statute arrives.