BREW
Prerequisites:
sudo apt install gcc build-essential
- Run these commands in your terminal to add Homebrew to your PATH:
echo >> /home/aidmin/.bashrc
echo 'eval "$(/home/linuxbrew/.linuxbrew/bin/brew shellenv bash)"' >> /home/aidmin/.bashrc
eval "$(/home/linuxbrew/.linuxbrew/bin/brew shellenv bash)"
- Install Homebrew's dependencies if you have sudo access:
For more information, see:
https://docs.brew.sh/Homebrew-on-Linux
- Run brew help to get started
- Further documentation:
https://docs.brew.sh
Start the llama.cpp server
brew install llama.cpp
llama-server -hf unsloth/Qwen3.6-27B-MTP-GGUF:
UD-Q4_K_XL
Configure Hermes
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default unsloth/Qwen3.6-27B-MTP-GGUF:
UD-Q4_K_XL
Run Hermes
Make a systemd service for llama-server
sudo nano /etc/systemd/system/llama-server.service
[Unit]
Description=Llama.cpp High-Performance LLM Server
After=network.target
[Service]
Type=simple
User=root
ExecStart=/usr/local/bin/llama-server -m /path/to/your/gemma-4-E4B.gguf -c 32768 --host 0.0.0.0 --port 8080
# The Magic: Auto-restart on ANY crash/freeze
Restart=always
RestartSec=5s
# Prevent systemd from giving up if it loops a couple of times
StartLimitIntervalSec=60s
StartLimitBurst=5
[Install]
WantedBy=multi-user.target
Save, Register, and Start It
# Reload the systemd manager to see the new service
sudo systemctl daemon-reload
# Enable it so it boots automatically if your LXC container or Proxmox host reboots
sudo systemctl enable llama-server
# Start it right now
sudo systemctl start llama-server
Change models script
systemd-llama-server.sh
#!/usr/bin/env bash
set -e
#---------- CONFIGURATION -----------------------------
MODEL_DIR="/home/aidmin/models_symlinks"
SERVICE_FILE="/etc/systemd/system/llama-server.service"
#------------------------------------------------------
# 1. Verify the GGUF folder exists
if [ ! -d "$MODEL_DIR" ]; then
echo "❌ Error: The directory $MODEL_DIR does not exist."
exit 1
fi
# 2. Collect all .gguf files into a list
cd "$MODEL_DIR"
GGUF_FILES=( *.gguf )
# Check if the directory is completely empty of models
if [ "${GGUF_FILES[0]}" == "*.gguf" ]; then
echo "❌ Error: No .gguf files found inside $MODEL_DIR"
exit 1
fi
echo "===================================================="
echo " LLAMA-SERVER INTERACTIVE MODEL SWAPPER "
echo "===================================================="
echo "Scanning: $MODEL_DIR"
echo "Please select a model number from the options below:"
echo "----------------------------------------------------"
# 3. Present the interactive menu prompt
PS3="Select a model (or enter 'q' to cancel): "
select SELECTED_MODEL in "${GGUF_FILES[@]}"; do
if [ "$REPLY" == "q" ] || [ "$REPLY" == "Q" ]; then
echo "Exiting without making changes."
exit 0
elif [ -n "$SELECTED_MODEL" ]; then
# Valid choice made, break out of menu loop
break
else
echo "❌ Invalid number choice. Please try again."
fi
done
# Create the full absolute path of the chosen model
FULL_MODEL_PATH="$MODEL_DIR/$SELECTED_MODEL"
echo "--------------------------------------------------"
echo "🎯 Selected: $SELECTED_MODEL"
echo "--------------------------------------------------"
# 4. Stop the active service
echo "🛑 Stopping llama-server.service to clear VRAM..."
sudo systemctl stop llama-server
# 5. Overwrite the ExecStart line inside the service file using sed
echo " ^=^s^} Rewriting systemd service configuration file..."
# This regex searches for any line starting with ExecStart= and swaps it out completely
#sudo sed -i "s|^ExecStart=.*|ExecStart=/home/linuxbrew/.linuxbrew/bin/llama-server -m $FULL_MODEL_PATH --host 0.0.0.0 --port 8080|" "$SERVICE_FILE"
Updated to improve performance.
sudo sed -i "s|^ExecStart=.*|ExecStart=/home/linuxbrew/.linuxbrew/bin/llama-server -m $FULL_MODEL_PATH --host 0.0.0.0 --port 8080 -ngl 99 --poll 0|" "$SERVICE_FILE"
# -ngl 99 (--n-gpu-layers 99): Offloads all model layers to your GPU, drastically improving inference speed compared to CPU-only execution.
# --poll 0: Prevents the CPU from maxing out at 100% usage during idle states while polling for incoming requests.
# 6. Tell systemd to register our file changes
sudo systemctl daemon-reload
# 7. Start the service back up
echo "⚡ Starting llama-server with the new model..."
sudo systemctl start llama-server
# 8. Service Verification Passthrough
echo "⏳ Waiting 3 seconds for VRAM allocation..."
sleep 3
echo "--------------------------------------------------"
echo "🔎 VERIFYING SYSTEM SERVICE HEALTH..."
echo "--------------------------------------------------"
if systemctl -q is-active llama-server; then
# Get the exact runtime and process information dynamically from systemd
ACTIVE_STATUS=$(systemctl show llama-server --property=ActiveState | cut -d= -f2)
SUB_STATUS=$(systemctl show llama-server --property=SubState | cut -d= -f2)
PID_NUM=$(systemctl show llama-server --property=MainPID | cut -d= -f2)
echo "🟢 STATUS: SUCCESS"
echo "📌 Service State : $ACTIVE_STATUS ($SUB_STATUS)"
echo "🆔 Main Process ID: $PID_NUM"
echo "🌐 Network Bind : http://127.0.0.1:8080/v1"
echo "--------------------------------------------------"
echo "✅ Done! Active model swapped and healthy."
else
echo "🔴 STATUS: FAILED"
echo "⚠️ Warning: llama-server failed to start cleanly."
echo "💡 Run 'journalctl -u llama-server -n 20' to check logs."
echo "--------------------------------------------------"
exit 1
fi
echo "=================================================="
Ways to Update Packages Installed with Brew
Homebrew handles both command-line tools (formulae) and graphical applications (casks) using a unified command set.
- Check what needs an update:Before upgrading, you can see which of your installed packages (including Llama or anything else) are outdated:
brew outdated
- Upgrade everything (Recommended):To update Homebrew itself and upgrade all outdated packages at once, run:
brew update && brew upgrade
- Upgrade a single specific package:If you only want to update one tool (for instance, just Llama) without touching anything else, specify its name:
brew upgrade llama
Important remote access capability
You can run the Hermes Web UI from a client by connecting via ssh. This is done by securely forward the port from the host (where Hermes is installed) to your local machine (client). First you need to start the Web UI service:
hermes dashboard --> command
Output:
HERMES_DASHBOARD_READY port=9119
Hermes Web UI → http://127.0.0.1:9119
Afther running the command, the server should be online. To stop it just ctrl + c.
Open another terminal and run:
ssh -L 9119:127.0.0.1:9119 your-username@<host-machine-lan-ip>
Open the browser on your client device and navigate to http://localhost:9119. It will securely tunnel straight to Hermes running on the host machine.
Is necessary to keep this terminal open, otherwise the connection will be closed. Logging out will close the tunnel and terminating the Hermes Web UI.
Hermes Available commands
+-------------------------------------------------------+
| (^_^)? Available Commands |
+-------------------------------------------------------+
── Session ──
/new - Start a new session (fresh session ID + history) (usage: /new [name])
/reset - Start a new session (fresh session ID + history) (alias for /new)
/clear - Clear screen and start a new session
/redraw - Force a full UI repaint (recovers from terminal drift)
/history - Show conversation history
/save - Save the current conversation
/retry - Retry the last message (resend to agent)
/undo - Back up N user turns and re-prompt (default 1) (usage: /undo [N])
/title - Set a title for the current session (usage: /title [name])
/handoff - Hand off this session to a messaging platform (Telegram, Discord, etc.) (usage: /handoff <platform>)
/branch - Branch the current session (explore a different path) (usage: /branch [name])
/fork - Branch the current session (explore a different path) (alias for /branch)
/compress - Compress conversation context (add 'here [N]' to keep recent N turns) (usage: /compress [here [N] | focus topic])
/rollback - List or restore filesystem checkpoints (usage: /rollback [number])
/snapshot - Create or restore state snapshots of Hermes config/state (usage: /snapshot [create|restore <id>|prune])
/snap - Create or restore state snapshots of Hermes config/state (alias for /snapshot)
/stop - Kill all running background processes
/background - Run a prompt in the background (usage: /background <prompt>)
/bg - Run a prompt in the background (alias for /background)
/btw - Run a prompt in the background (alias for /background)
/agents - Show active agents and running tasks
/tasks - Show active agents and running tasks (alias for /agents)
/queue - Queue a prompt for the next turn (doesn't interrupt) (usage: /queue <prompt>)
/q - Queue a prompt for the next turn (doesn't interrupt) (alias for /queue)
/steer - Inject a message after the next tool call without interrupting (usage: /steer <prompt>)
/goal - Set a standing goal Hermes works on across turns until achieved (usage: /goal [text | pause | resume | clear | status])
/subgoal - Add or manage extra criteria on the active goal (usage: /subgoal [text | remove N | clear])
/status - Show session info
/resume - Resume a previously-named session (usage: /resume [name])
/sessions - Browse and resume previous sessions
── Info ──
/whoami - Show your slash command access (admin / user)
/profile - Show active profile name and home directory
/gquota - Show Google Gemini Code Assist quota usage
/help - Show available commands
/usage - Show token usage and rate limits for the current session
/insights - Show usage insights and analytics (usage: /insights [days])
/platforms - Show gateway/messaging platform status
/gateway - Show gateway/messaging platform status (alias for /platforms)
/copy - Copy the last assistant response to clipboard (usage: /copy [number])
/paste - Attach clipboard image from your clipboard
/image - Attach a local image file for your next prompt (usage: /image <path>)
/update - Update Hermes Agent to the latest version
/debug - Upload debug report (system info + logs) and get shareable links
── Configuration ──
/config - Show current configuration
/model - Switch model for this session (usage: /model [model] [--provider name] [--global] [--refresh])
/codex-runtime - Toggle codex app-server runtime for OpenAI/Codex models (usage: /codex-runtime [auto|codex_app_server])
/codex_runtime - Toggle codex app-server runtime for OpenAI/Codex models (alias for /codex-runtime)
/personality - Set a predefined personality (usage: /personality [name])
/statusbar - Toggle the context/model status bar
/sb - Toggle the context/model status bar (alias for /statusbar)
/verbose - Cycle tool progress display: off -> new -> all -> verbose
/footer - Toggle gateway runtime-metadata footer on final replies (usage: /footer [on|off|status])
/yolo - Toggle YOLO mode (skip all dangerous command approvals)
/reasoning - Manage reasoning effort and display (usage: /reasoning [level|show|hide])
/skin - Show or change the display skin/theme (usage: /skin [name])
/indicator - Pick the TUI busy-indicator style (usage: /indicator [kaomoji|emoji|unicode|ascii])
/voice - Toggle voice mode (usage: /voice [on|off|tts|status])
/busy - Control what Enter does while Hermes is working (usage: /busy [queue|steer|interrupt|status])
── Tools & Skills ──
/tools - Manage tools: /tools [list|disable|enable] [name...] (usage: /tools [list|disable|enable] [name...])
/toolsets - List available toolsets
/skills - Search, install, inspect, or manage skills
/bundles - List skill bundles (aliases /<name> for multiple skills)
/cron - Manage scheduled tasks (usage: /cron [subcommand])
/curator - Background skill maintenance (status, run, pin, archive, list-archived) (usage: /curator [subcommand])
/kanban - Multi-profile collaboration board (tasks, links, comments) (usage: /kanban [subcommand])
/reload - Reload .env variables into the running session
/reload-mcp - Reload MCP servers from config
/reload_mcp - Reload MCP servers from config (alias for /reload-mcp)
/reload-skills - Re-scan ~/.hermes/skills/ for newly installed or removed skills
/reload_skills - Re-scan ~/.hermes/skills/ for newly installed or removed skills (alias for /reload-skills)
/browser - Connect browser tools to your live Chromium-family browser via CDP (usage: /browser [connect|disconnect|status])
/plugins - List installed plugins and their status
── Exit ──
/quit - Exit the CLI (use --delete to also remove session history) (usage: /quit [--delete])
/exit - Exit the CLI (use --delete to also remove session history) (alias for /quit)
⚡ Skill Commands (70 installed):
/airtable - Airtable REST API via curl. Records CRUD, filters, upserts.
/architecture-diagram - Dark-themed SVG architecture/cloud/infra diagrams as HTML.
/arxiv - Search arXiv papers by keyword, author, category, or ID.
/ascii-art - ASCII art: pyfiglet, cowsay, boxes, image-to-ascii.
/ascii-video - ASCII video: convert video/audio to colored ASCII MP4/GIF.
/audiocraft-audio-generation - AudioCraft: MusicGen text-to-music, AudioGen text-to-sound.
/baoyu-infographic - Infographics: 21 layouts x 21 styles (信息图, 可视化).
/blogwatcher - Monitor blogs and RSS/Atom feeds via blogwatcher-cli tool.
/claude-code - Delegate coding to Claude Code CLI (features, PRs).
/claude-design - Design one-off HTML artifacts (landing, deck, prototype).
/codebase-inspection - Inspect codebases w/ pygount: LOC, languages, ratios.
/codex - Delegate coding to OpenAI Codex CLI (features, PRs).
/comfyui - Generate images, video, and audio with ComfyUI — install, launch, manage nodes/models, run workflows with parameter injection. Uses the official comfy-cli for lifecycle and
direct REST/WebSocket API for execution.
/design-md - Author/validate/export Google's DESIGN.md token spec files.
/dogfood - Exploratory QA of web apps: find bugs, evidence, reports.
/evaluating-llms-harness - lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).
/excalidraw - Hand-drawn Excalidraw JSON diagrams (arch, flow, seq).
/gif-search - Search/download GIFs from Tenor via curl + jq.
/github-auth - GitHub auth setup: HTTPS tokens, SSH keys, gh CLI login.
/github-code-review - Review PRs: diffs, inline comments via gh or REST.
/github-issues - Create, triage, label, assign GitHub issues via gh or REST.
/github-pr-workflow - GitHub PR lifecycle: branch, commit, open, CI, merge.
/github-repo-management - Clone/create/fork repos; manage remotes, releases.
/godmode - Jailbreak LLMs: Parseltongue, GODMODE, ULTRAPLINIAN.
/google-workspace - Gmail, Calendar, Drive, Docs, Sheets via gws CLI or Python.
/heartmula - HeartMuLa: Suno-like song generation from lyrics + tags.
/hermes-agent - Configure, extend, or contribute to Hermes Agent.
/hermes-agent-skill-authoring - Author in-repo SKILL.md: frontmatter, validator, structure.
/himalaya - Himalaya CLI: IMAP/SMTP email from terminal.
/huggingface-hub - HuggingFace hf CLI: search/download/upload models, datasets.
/humanizer - Humanize text: strip AI-isms and add real voice.
/jupyter-live-kernel - Iterative Python via live Jupyter kernel (hamelnb).
/kanban-orchestrator - Decomposition playbook + anti-temptation rules for an orchestrator profile routing work through Kanban. The "don't do the work yourself" rule and the basic lifecycle are
auto-injected into every kanban worker's system prompt; this skill is the deeper playbook when you're specifically playing the orchestrator role.
/kanban-worker - Pitfalls, examples, and edge cases for Hermes Kanban workers. The lifecycle itself is auto-injected into every worker's system prompt as KANBAN_GUIDANCE (from
agent/prompt_builder.py); this skill is what you load when you want deeper detail on specific scenarios.
/llama-cpp - llama.cpp local GGUF inference + HF Hub model discovery.
/llm-wiki - Karpathy's LLM Wiki: build/query interlinked markdown KB.
/manim-video - Manim CE animations: 3Blue1Brown math/algo videos.
/maps - Geocode, POIs, routes, timezones via OpenStreetMap/OSRM.
/nano-pdf - Edit PDF text/typos/titles via nano-pdf CLI (NL prompts).
/network-system-analysis - Performs a comprehensive network and system diagnostics check. This skill is used to audit active processes (via lsof), check listening ports (via ss), review configuration
(via config.yaml), and verify external connections (like GitHub/Tirith). It outputs a structured report summarizing findings, security posture, and recommended actions.
/node-inspect-debugger - Debug Node.js via --inspect + Chrome DevTools Protocol CLI.
/notion - Notion API + ntn CLI: pages, databases, markdown, Workers.
/obliteratus - OBLITERATUS: abliterate LLM refusals (diff-in-means).
/obsidian - Read, search, create, and edit notes in the Obsidian vault.
/ocr-and-documents - Extract text from PDFs/scans (pymupdf, marker-pdf).
/opencode - Delegate coding to OpenCode CLI (features, PR review).
/openhue - Control Philips Hue lights, scenes, rooms via OpenHue CLI.
/p5js - p5.js sketches: gen art, shaders, interactive, 3D.
/plan - Plan mode: write an actionable markdown plan to .hermes/plans/, no execution. Bite-sized tasks, exact paths, complete code.
/polymarket - Query Polymarket: markets, prices, orderbooks, history.
/popular-web-designs - 54 real design systems (Stripe, Linear, Vercel) as HTML/CSS.
/powerpoint - Create, read, edit .pptx decks, slides, notes, templates.
/pretext - Use when building creative browser demos with @chenglou/pretext — DOM-free text layout for ASCII art, typographic flow around obstacles, text-as-geometry games, kinetic
typography, and text-powered generative art. Produces single-file HTML demos by default.
/python-debugpy - Debug Python: pdb REPL + debugpy remote (DAP).
/requesting-code-review - Pre-commit review: security scan, quality gates, auto-fix.
/research-paper-writing - Write ML papers for NeurIPS/ICML/ICLR: design→submit.
/segment-anything-model - SAM: zero-shot image segmentation via points, boxes, masks.
/serving-llms-vllm - vLLM: high-throughput LLM serving, OpenAI API, quantization.
/sketch - Throwaway HTML mockups: 2-3 design variants to compare.
/songsee - Audio spectrograms/features (mel, chroma, MFCC) via CLI.
/songwriting-and-ai-music - Songwriting craft and Suno AI music prompts.
/spike - Throwaway experiments to validate an idea before build.
/systematic-debugging - 4-phase root cause debugging: understand bugs before fixing.
/teams-meeting-pipeline - Operate the Teams meeting summary pipeline via Hermes CLI — summarize meetings, inspect pipeline status, replay jobs, manage Microsoft Graph subscriptions.
/test-driven-development - TDD: enforce RED-GREEN-REFACTOR, tests before code.
/touchdesigner-mcp - Control a running TouchDesigner instance via twozero MCP — create operators, set parameters, wire connections, execute Python, build real-time visuals. 36 native tools.
/weights-and-biases - W&B: log ML experiments, sweeps, model registry, dashboards.
/xurl - X/Twitter via xurl CLI: post, search, DM, media, v2 API.
/youtube-content - YouTube transcripts to summaries, threads, blogs.
/yuanbao - Yuanbao (元宝) groups: @mention users, query info/members.