Ultra memory-frugal layer-by-layer weight streaming, dynamic Wi-Fi multi-device tensor mesh sharding, sovereign multi-agent workflow DAGs, and native multimodal vision/voice reasoning.
Type commands below or click the quick action chips to test the engine CLI interactively.
A unified suite featuring a Web Workspace UI, dynamic P2P Wi-Fi mesh topology, and voice/vision pipelines.
Select your GPU VRAM size and desired LLM model to instantly calculate speed, RAM allocation, and execution mode.
Real-world benchmarks comparing token generation speeds and VRAM footprint across consumer GPUs.
Answers to common questions about hardware compatibility, layer streaming, and privacy.
NEW BRAIN uses an ultra-frugal layer-by-layer weight streaming engine. Instead of loading all 40GB+ of model weights into GPU VRAM at once, it iteratively streams layers from system RAM/NVMe into GPU memory during execution, executing layer calculations and flushing memory without triggering CUDA Out-of-Memory errors.
Running `/share` or scanning the QR code connects nearby laptops, PCs, and Macs over local Wi-Fi or LAN. NEW BRAIN automatically splits model layers across connected devices, pooling their combined VRAM and RAM into a single unified supercomputing cluster.
Yes. NEW BRAIN runs 100% locally on your machine. Zero telemetry, zero external server calls, and zero data logging outside your local network.
Absolutely. NEW BRAIN runs a standard OpenAI-compatible REST server on `http://localhost:8080/v1/chat/completions` and Anthropic Claude endpoint compatibility. Any tool supporting OpenAI format can connect seamlessly.
See how NEW BRAIN compares against standard local AI tools and expensive cloud APIs.
| Feature Capability | NEW BRAIN Engine | Standard Local LLM Tools (Ollama / LM Studio) | Cloud LLM APIs (OpenAI / Anthropic) |
|---|---|---|---|
| 80+ AI Models Compatible | ✓ Yes (DeepSeek R1, Llama, Qwen) | ✓ Partial | ✗ Proprietary Only |
| Run 70B Models on 4GB GPU | ✓ Yes (Layer Streaming) | ✗ Requires $2,000+ GPUs | ✗ Pay-Per-Token Cloud |
| Wi-Fi Multi-Device Mesh Sharding | ✓ Yes (P2P Tensor Mesh) | ✗ Single Machine Only | ✗ Centralized Cloud |
| Autonomous Multi-Agent DAGs | ✓ Built-in Teams & Schema Guards | ✗ Basic Single Scripts | ✗ Custom API Glue Code |
| Native Multimodal (Voice + Vision) | ✓ Audio, Video Frames, Vision & Speech | ✗ Text Only / Limited | ✓ High-Cost API Tiers |
| 100% Local Privacy & Sovereignty | ✓ 100% Offline & Private | ✓ Local | ✗ Third-Party Data Exposure |
| OpenAI / Anthropic REST API Endpoint | ✓ Yes (Port 8080) | ✓ Yes | ✓ Native |
Built from the ground up for modern hardware constraints, mobile connectivity, and distributed edge AI.
Native out-of-the-box support for DeepSeek R1/V3, Llama 3.3/3.2, Qwen 2.5/Coder, Mistral/Mixtral, Gemma 2, Phi-4, Whisper, and 80+ open-source models.
Iteratively streams model tensor layers from disk/RAM into GPU memory during execution, allowing 70B parameter models to run flawlessly on consumer 4GB VRAM graphics cards.
Automatic 4-bit, 2-bit, FP8, and GGUF model weight quantization engine. Compresses raw HuggingFace weights instantly on load to fit constrained VRAM budgets.
Shard model tensor layers dynamically across idle laptops, PCs, and Macs over local Wi-Fi or LAN networks without expensive interconnects or server racks.
Instantly pair smartphones (iOS / Android) to your local AI engine via QR code scanning. Full mobile responsive Web UI (`http://
Orchestrate autonomous agent teams with stateful DAG workflows, strict response schema validation guards, and human-in-the-loop approval management.
Ingest and reason over text, live microphone speech audio, real-time video frame extraction, images, and webcam streams directly inside the CLI and Web workspace.
Automatic 404 detection and backup key rotation failover for hybrid local and cloud model inference pipelines.
Upcoming sovereign operating system and security frameworks bringing bare-metal AI boot layers and Red/Blue/Grey/Black team agents directly to your network.
A lightweight, bare-metal sovereign AI operating system with integrated P2P mesh kernel drivers, real-time hardware weight streaming, native autonomous agent desktop workspace, and zero-latency voice/vision system services.
⚡ In Architecture & Kernel Prototyping Phase (v3.0 Roadmap)Self-directed Red Team AI agents for vulnerability auditing, penetration scenario simulation, posture evaluation, and automated exploit payload validation.
⚡ In Active Development (v2.0)Real-time network anomaly detection, automated firewall rule generation, incident response DAGs, and local log triage orchestration.
⚡ In Active Development (v2.0)Continuous compliance verification, system access mapping, policy enforcement, and multi-tier security risk analysis across your P2P mesh node topology.
⚡ In Active Development (v2.0)Advanced adversarial AI agents engineered to simulate zero-day threats, stress-test air-gapped networks, and evaluate perimeter defense resilience.
⚡ Research & Testing PhaseLocal offline RTSP/IP camera discovery, default credential auditing, stream integrity checking, and AI vision anomaly detection without external cloud dependencies.
⚡ Planned ModuleIntegrated penetration testing framework supporting automated port discovery, vulnerability report generation, and remediation playbooks.
⚡ Planned ModuleTell us what features, optimizations, or agent capabilities you want to see in the next version of NEW BRAIN!