Wood Chen

llmfit: Find Out Which Local LLMs Your Computer Can Run with One Command

0 comments3 views569 words

This post was translated from Chinese by AI. If anything reads oddly, the Chinese original is authoritative. 中文原文

If you've tried running large language models locally, you've probably hit the same wall: you find a model you like, download a dozen or so GB, then run out of VRAM when loading it—or barely get it running at two or three tokens per second. llmfit aims to help you make that call before downloading: it reads your hardware specs and tells you which models you can run and roughly how fast they'll be. It has climbed the Rust trending list over the past couple of days.

llmfit demo: searching models, simulating different hardware, and planning deployments

What it solves

llmfit is a terminal tool. On startup, it detects your CPU core count, RAM, discrete/integrated GPUs, VRAM, and unified memory architectures such as Apple Silicon (supporting NVIDIA CUDA, AMD ROCm, Intel OneAPI, and Apple Silicon). It then scores each model in its built-in catalog across four dimensions: whether it fits in memory, estimated speed, quality, and context length. Results are ranked by fit, so you don't have to calculate "how much VRAM a 7B model at Q4 needs" yourself.

A few features I find useful:

  • Accounts for quantization formats (GGUF, AWQ, GPTQ, EXL2), showing memory usage and speed for different quantizations of the same model
  • Recognizes MoE models and estimates speed based on active parameters rather than total parameters, making it more reliable than many "VRAM calculators"
  • Supports multiple GPUs
  • Bases speed estimates on a memory bandwidth model; you can inspect the assumptions behind each number with llmfit info
  • Integrates with local runtimes such as Ollama, llama.cpp, MLX, Docker Model Runner, and LM Studio
  • Also includes a Web dashboard and REST endpoints (/api/v1/system, /api/v1/models) that scripts and agents can call directly

The recently added benchmark feature is interesting too: you can download a model locally, start a service, and measure actual tok/s. The results replace the estimates in the table, and you can submit a PR directly from the TUI to contribute them to the community. Others with the same hardware will then see real measurements.

How to use it

There are plenty of installation options:


scoop install llmfit

brew install AlexsJones/llmfit/llmfit

uv tool install -U llmfit
uvx llmfit        # Run directly without installing

Common commands:

llmfit                       # Interactive TUI: your hardware + all models, ranked by fit
llmfit fit                   # Classic table output
llmfit recommend --json      # Output recommendations as JSON for scripts/agents
llmfit info "<模型名>"        # Fit analysis and estimation assumptions for a single model
llmfit bench                 # Measure tok/s and time to first token on a running service
llmfit serve --port 8787     # Start the Web interface and API

In the TUI, / filters by name, family, or quantization, and h shows help. Machines without a GPU work too—it will make recommendations based on your CPU and RAM.

You can also run it as a persistent dashboard on a server. An official Docker image is available:

docker run -d -p 8787:8787 ghcr.io/alexsjones/llmfit serve

Who it's for, and what to watch out for

It's useful if you're choosing a local model or planning which model each machine in a fleet should run. Combine recommend --json with jq, and you can make it a step in your deployment workflow.

Keep in mind that, until you run a benchmark, the speed figures are estimates, not measurements. The README also notes that a similar tool, llm-checker, takes the approach of "actually running the model through Ollama." That gets closer to real-world performance but requires installing Ollama first. llmfit estimates from hardware specs upfront, trading accuracy for convenience. You can use both: shortlist models with llmfit, then benchmark the finalists.

Project details

  • Language: Rust
  • License:MIT
  • Star: approximately 3.7 tens of thousands as of 2026-10-01
  • Latest version: v1.1.16(2026-09-19)
  • GitHub:https://github.com/AlexsJones/llmfit

Originally published on the SunAI forum. Versions, star counts, and relative dates in this article reflect the time of the original post.

Last updated 2026-10-09

Related posts

Comments 0