Skip to content

MCP

Local models with LM Studio

Every assistant in this section runs a frontier model in the cloud, and most of them are paid. LM Studio is the way to try RhinoArtisan MCP with a model running on your own machine: free, offline, and nothing leaves your computer.

Frontier models vs local models

A frontier model is one of the largest, most capable models that exist — Anthropic’s Claude, OpenAI’s GPT, Google’s Gemini. They are trained and run by the big AI labs on hardware no desktop can match, and they are the models that make MCP shine: they follow a long list of tools, chain several calls to complete a task, recover from errors and understand jewelry vocabulary without being told.

A local model is a smaller, open-weights model — the Llama, Qwen, Mistral, Gemma or gpt-oss families — compressed (quantized) so it fits on a consumer computer. Fewer parameters and less precision mean less capability: a local model may pick the wrong tool, invent an argument, lose the thread after a few steps, or give up on a long instruction. The best ones handle simple, one-step requests well. None of them matches a frontier model.

Rule of thumb: what a frontier model does in one prompt, a local model may need several — or may never do.

Hardware

A local model runs entirely on your machine, and the bigger the model, the better it follows tools and the more memory it needs. The model has to fit in memory: unified memory on Apple Silicon, GPU memory (VRAM) on Windows. If it doesn’t fit, LM Studio falls back to the CPU and the model becomes too slow to be useful.

Rough sizes at 4-bit quantization, the usual compromise:

Model sizeMemory neededWhat to expect with RhinoArtisan MCP
7–9B parameters~5–6 GBSimple one-step tasks, frequent mistakes
14B~9–10 GBUsable for short conversations
30B and up20 GB or moreThe closest a local model gets to a frontier one

Where to check:

  • LM Studio’s system requirements — supported chips and operating systems. In short: Apple Silicon on macOS 14 or newer (Intel Macs are not supported), Windows x64 with AVX2 or ARM, 16 GB of RAM recommended, at least 4 GB of dedicated VRAM on Windows.
  • The model page inside LM Studio — every download shows its size and whether it fits your machine before you download it.
  • Context length. RhinoArtisan exposes 205 tools, and their definitions alone take a big slice of the model’s context window. Load the model with the largest context your memory allows — 32k tokens or more — or the conversation overflows after a couple of messages. LM Studio’s own docs warn that MCP servers built for Claude or ChatGPT “may quickly bog down your local model”.

Setup

You need LM Studio 0.3.17 or newer, a model with the hammer badge (native tool use — LM Studio shows it in the model browser), and Rhino running with RhinoArtisan 7.

  1. Install LM Studio and download a model with the hammer badge. Pick the largest one that fits your machine.

  2. Add RhinoArtisan with the one-click install: Add to LM Studio

    Or by hand: open the Program tab in the right sidebar, choose Install → Edit mcp.json, and add:

    {
      "mcpServers": {
        "artisan": {
          "url": "http://127.0.0.1:9280/mcp"
        }
      }
    }
  3. Load the model with a large context length, and turn the artisan server on for your chat from the Program tab.

Using a different port? If Rhino printed another address when you ran ArtisanMcpStart Port=…, use that one instead of http://127.0.0.1:9280/mcp. See Changing the port.

Try it

With Rhino open, ask:

“Create a round 1 ct diamond.”

LM Studio asks you to confirm the tool call the first time — you can allow it once or always, and manage those permissions later under App Settings → Tools & Integrations. The stone appears in your Rhino viewport.

What to expect

  • The model never calls a tool. It has no native tool support, or the context is too small for the tool list. Try a model with the hammer badge and a larger context.
  • It calls the wrong tool or invents arguments. That is the model’s capability, not a RhinoArtisan issue. A bigger model helps; if a bigger one doesn’t fit, you have reached your hardware’s ceiling.
  • Everything is slow. The model doesn’t fit in GPU memory. Use a smaller model or a smaller quantization.
  • A tool returns an error with Rhino open. Check Troubleshooting first. If the same request works from Claude Desktop and fails from a local model, the difference is the model.