ECC
— Open agent harness for coding workflows, GitHub automation, and securityOpen agent harness for coding workflows, GitHub automation, and security

High-throughput, memory-efficient LLM inference and serving
vLLM is an inference and serving engine for large language models (LLMs). It is intended for developers and teams deploying models across different hardware platforms.
The engine provides a drop-in OpenAI-compatible API and supports a range of hardware, including NVIDIA CUDA GPUs, AMD ROCm GPUs, AWS Neuron accelerators, Google Cloud TPUs, Apple Silicon, and others. PagedAttention, advanced scheduling, and continuous batching are used to maximize throughput and GPU utilization.
vLLM can be installed with Python or Docker. The quick-start instructions require Python 3.10 or newer, with Python 3.12 or newer recommended; stable and nightly builds are available.
Used it? Write the first review.
Open agent harness for coding workflows, GitHub automation, and security
Open-source self-hosted AI agent with persistent memory
A modular, plugin-based framework for building and running agents
Context API for searching, scraping, and interacting with the web

An AI coding agent for terminal, IDE, web, and Slack
Self-hosted interface for connecting and extending AI models