Tabby
Self-hosted AI coding assistant
What is Tabby?
Tabby is an open-source, self-hosted AI coding assistant that provides code completion and chat inside your editor. Your code never leaves your infrastructure, which makes it a practical choice for teams with strict code privacy requirements.
Best for
Teams that cannot send source code to third-party AI services
Why choose Tabby
Tabby is for developers who want code completion without shipping their source code to a hosted service. It runs on your own machine or server — a consumer GPU is enough for a usable setup — and plugs into VS Code, JetBrains and Vim, providing inline completion and a chat that can answer questions about your own repository. For anyone working on code they are contractually or ethically unable to paste into a cloud assistant, this is the practical path to the same convenience.
Replaces
- GitHub Copilot
- Codeium
- Amazon CodeWhisperer
Key features
- Code completion and inline chat
- Editor plugins for VS Code, JetBrains, Vim
- Runs on consumer GPUs
- Answer engine over your codebase
What to watch out for
Model quality is the constraint, and it is a real one: a self-hosted model on consumer hardware will not match the best hosted assistants on complex reasoning or unfamiliar libraries, and completion quality varies a lot with language and codebase style. Setting it up properly involves running a model server and tuning it against your hardware, which is a weekend of work. The project's priorities have shifted over time, so check that the features you need are actively maintained before building a workflow on them.
How to deploy
- Docker
- Binary
Getting started
Check the hardware requirements honestly before installing — a machine without a capable GPU can run a small model but the experience will be modest. Get the server running and complete a code completion in the editor before configuring anything advanced, so you can tell a model problem apart from a plugin problem. Point the answer engine at a single repository first and evaluate the results on questions you already know the answers to. If completion quality disappoints, adjust the model before you adjust the editor settings.
Typical setup
It runs on a developer's workstation or a small GPU server on the local network, with the editor plugin pointing at it. Teams that care about code confidentiality run one instance on an internal machine and have everyone's editors connect to it, so completions are generated without source code leaving the building. The setup that works best is a machine with a capable GPU, a model sized to that hardware, and realistic expectations about which languages it handles well.
Who should look elsewhere
Not a substitute for a frontier hosted assistant if your goal is the best possible completions on unfamiliar code. Self-hosted models on consumer hardware trail significantly on difficult reasoning, and setting up a model server is real work. Anyone without a capable GPU, or without the appetite to tune a model, will get more value from a hosted tool.
Project health
- GitHub stars: 33,885
- Last code push: 2026-06-30
- Open issues: 343
- Status: actively developed
Figures pulled from the GitHub API and refreshed periodically.