Open-source AI coding assistant for VS Code. Code completion, chat, edits and reviews with local or hosted models. Your models, your infrastructure.
-
Updated
Oct 2, 2026 - TypeScript
Open-source AI coding assistant for VS Code. Code completion, chat, edits and reviews with local or hosted models. Your models, your infrastructure.
A universal OpenCode plugin for dynamic model discovery with flexible configuration for OpenAI-compatible providers.
Self-hosted Qwen3.8-27B (FP8) inference with vLLM, KServe and Envoy AI Gateway on RTX 6000 PRo or 2× RTX 4080 Super
Benchmark how well any model performs real-world, long-horizon coding tasks across coding agents (harnesses) -- Claude Code and pi. Run models three ways -- Anthropic on Amazon Bedrock, open-weight on Bedrock via a LiteLLM proxy, or self-hosted on EC2 with vLLM -- score them with an LLM judge, and plot the cost/quality Pareto frontier.
Terraform setup for deploying a private coding LLM on Vast.ai with vLLM, Qwen3 Coder, and OpenCode.
Free ChatGPT API & Free LLM API Key Alternative. 100% local drop-in OpenAI replacement supporting System Prompts, Tool Calling, and Think Mode for AI Agents.
Enterprise-grade Sovereign AI Stack optimized for NVIDIA Blackwell (sm_120) & vLLM. Features 256K context window, 5.8k tok/s prefill, and integrated observability via Langfuse.
The control plane for self-hosted AI inference. Warm-state GPU routing, multi-runtime orchestration across Ollama, vLLM, llama.cpp, TGI and MLX . Single Go binary. Apache-2.0.
A platform-agnostic format + method for sizing local-LLM hardware from your real agent sessions, calibrated with capability-oracle runs.
Windows utility for using Codex Desktop GUI with Ollama and other local/self-hosted LLM backends via profiles.
Alternatively prompt, an LLM-based plugin for the IntelliJ product family. Answer questions and generate code with self-hosted models.
Production-grade multimodal RAG assistant using open-source LLMs and vector databases.
Air-gapped pre-deployment network change validation against a real containerlab digital twin, with sealed PCI/SOC2/NIST evidence. Zero egress.
GLM-5.2 744B at 4-bit on Modal 4x H200 via vLLM, plus a static streaming chat UI.
Self-hosted AI coding platform — provisions GPU compute on Vast.ai and deploys open-weight LLMs via llama.cpp as a bring-your-own-key backend for GitHub Copilot (also auto-configures Continue.dev/Cline).
A university-scale LLM serving platform in miniature: vLLM on cloud GPU, LiteLLM gateway with per-faculty governance, Prometheus/Grafana SLOs, k3d/ArgoCD GitOps. All numbers measured, all failures documented.
Self-hosted vLLM inference stack with an OpenAI-compatible API, Docker Compose templates, Caddy reverse proxy, and NVIDIA GPU thermal guard.
AI-powered herbal remedies chatbot based on "The Little Handbook of Natural Remedies" by Michael Martin. RAG system for natural medicine research.
Protocore — the agent loop, and nothing else. A protocol-first ReAct runtime for LLM agents in Python 3.12+: context budget, tool surface, compaction, snapshot and resume. No database driver, no HTTP endpoint — everything outward is a Protocol you implement. MPL-2.0.
Self-hosted app for comparing local LLMs — live chat, A/B model diffing, and a parameter-sweep eval harness with LLM-judge scoring.
To associate your repository with the self-hosted-llm topic, visit your repo's landing page and select "manage topics."