Downloading LM Studio and running the Prompt

Basics

Most university licences are confidential, and most contain a clause restricting disclosure of the agreement to third parties. Sending the licence to a cloud-hosted LLM (Claude, Gemini , etc.) is a disclosure to a third party, whatever that provider’s security posture or retention policy happens to be. Consumer accounts on some services also allow uploaded content to be retained and used for model improvement depending on account settings.

You may be able to satisfy yourself that a particular enterprise agreement, with a data processing agreement in place, is compatible with your obligations. That analysis takes time, has to be redone whenever the provider’s terms change, and is a conversation you probably do not want to have with your TTO in the middle of a negotiation. Running the model on your own machine removes the question entirely since the licence never leaves your computer.

The instructions below enable you to download and run a local model and then block access to the internet either by simply disconnecting all network connections (e.g. run in Airplane mode) or use a Firewall to block the model making outbound connections. Note the order: install LM Studio and download the model while you are online, then disconnect. Once the model is on your machine, no network connection is needed to run it, and disconnecting is a simple way to prove to yourself, and to your TTO if asked, that nothing left the computer

Several tools let you run open-weight models locally; the examples below all use LM Studio because it has the friendliest document-handling interface we have found.

Download and Installation

LM Studio is available at https://lmstudio.ai/download

Click the Download LM Studio button (not Bionic) and follow the installation instructions.

For those Arm fans out there, there is also an Arm (Snapdragon) variant, and the models have also been tested on those.

Models

Models are updated regularly and for new ones look in Model Search. The prompt examples below use the qwen/qwen3.5-9b model, which is a less than 10 Gbyte sized model that should run within most desktop environments.

If you have a desktop with more memory, then feel free to choose a larger sized model, but note that we have tested the qwen/qwen3.5-9b model against multiple different licences, and multiple different models and it’s hard to beat.

Models are several gigabytes each. On Windows, if your C: drive is short of space, go to the My Models tab (folder icon), click Settings and change the Models Directory to a folder on another drive. On macOS the equivalent setting points by default at a folder in your home directory; change it in the same place if you would rather keep models on an external volume. Apple Silicon Macs run these models particularly well.

Settings

Default settings are fine except for one: context length. LM Studio loads models with a modest default context, which is not enough to hold a full licence plus the Prompt. Go to Model Defaults and raise the Default Context Length; 32,768 tokens is a sensible starting point for a typical licence, and this model supports far more if your machine can take it. Be aware that context length is the main driver of memory use. If the model fails to load, or the machine starts swapping to disk, reduce the context length or the GPU offload before assuming the model is too large.

Helping in negotiations

The Prompt file incorporates the key negotiation points in the Framework.

Upload your chosen Licence and the Prompt into the “Send a message to the model …” window either by dragging and dropping them, or using the + entry in the window, and then just hit return. You should then see something very similar to the message below. If you see anything other than “inject-full-context“, stop and do not rely on the output. It means the licence and prompt together exceed the context, and the model is working from extracts rather than the whole agreement. It will still produce a confident-looking answer, but it will be wrong. Fix it by raising the Default Context Length (see Settings above), then start a new chat and re-attach the files. If you cannot raise the context far enough on your hardware, use a smaller model rather than accepting a retrieval run.

16 GB of RAM is the practical minimum; 32 GB is comfortable and lets you use a longer context. A discrete GPU or Apple Silicon makes a substantial difference to speed — on an M-series Mac a full licence review takes a few minutes, while on a CPU-only laptop it can take 20–30 minutes. Whichever you have, change your power plan so the machine does not sleep mid-run, and close other memory-hungry applications.

Once it runs, you will have a complete set of differences from the SDL and best practice, and more importantly, a set of recommendations, justifications and specific actions to take to your TTO to make the negotiations more fruitful and hopefully quicker. It’s AI produced, so do check the output before you use it. Local models are good at spotting where a licence departs from the SDL, but they do misread defined terms, miss cross-references between clauses, and occasionally cite clause numbers that do not exist. Treat every finding as a pointer back to the licence text: open the clause, read it, and satisfy yourself the point is real before you put it in front of your TTO. Nothing produced by this process is legal advice, and it is not a substitute for taking advice on your own agreement.

You can then ask further clarification questions of the model such as “If we do a funding raise of £500K in 6 months and have zero revenue, then how do the licences compare in terms of cash outlay by the Company”, etc. This helps you work through different scenarios.