noonghunna/club-3090:README 來源編輯指南
根據 README、倉庫資料與授權整理 noonghunna/club-3090 的安裝與核驗路徑。
專案定位
noonghunna/club-3090 的 README 將專案描述為「Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ikllama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.」。本文只整理倉庫可直接核對的內容,不把 star、Fork 或宣傳語當成品質證明。README 在「club-3090」下寫到:Recipes for serving LLMs locally on RTX 3090s. Multi-engine (vLLM, llama.cpp, ikllama), multi-model, model-agnostic by design.。這說明的是專案邊界,不是已完成的生產驗證。
適用場景
從 README 的「TL;DR , what this is」與相關條目,可以先判斷它是否處理你的實際問題:🏎 vLLM dual = max throughput. Up to 127 TPS code (DFlash) or 4 concurrent streams @ 262K (turbo). Full feature stack (vision · tools · MTP · streaming).。若需求不同,不應只因專案熱度就採用。本文保留原始專案名、命令與元件名,方便回到一手來源核對。 README 另外列出一項可核對的資訊:Two complementary routes , pick by what your workload breaks on:。這類原文條目可用來設計試跑步驟,但不能取代實際環境測試。
運作方式
README 將運作方式分散在「club-3090」等段落。可確認的線索包括:> 🎯 4090 or 5090 owner? The composes run cross-rig , contributors have benched both with measured numbers: Can I use a 4090? → + cross-rig benchmark rows live in the FAQ.。本文不把未寫出的架構、效能或安全邊界補成結論;真正的執行鏈仍要配合目錄、設定檔與版本標籤檢查。
安裝與第一次執行
第一次安裝應從 README 指出的入口開始。目前可核對的命令是: # 1. Clone the repo git clone https://github.com/noonghunna/club-3090.git cd club-3090 # Profile compatibility tooling requires PyYAML. Ubuntu LTS usually has it via # python3-yaml; otherwise run: python3 -m pip install pyyaml # 2. Pick/download + SHA-verify the model (interactive hardware-aware picker) # (asks you which model, then where to put model weights , pick in-repo # default, ~/models, or a custom path on a different drive. To skip prompts: # `export MODEL_DIR=/path/to/mode 如果倉庫沒有命令,本文不會自行編造步驟,而是建議先閱讀「Quick start」,確認系統依賴、預設埠與首次初始化。
設定與日常使用
日常使用取決於專案文件。README 的「club-3090」段落提到:> 🎨 Want image generation too? The Image Studio bundle runs Ideogram-4 image gen + a chat model + Open WebUI together on two GPUs , one command: bash scripts/setup-image-studio.sh.。設定檔、環境變數、權限與資料目錄只在來源明確時才會記錄;沒有寫出的預設值,應在測試環境驗證並保留回滾副本。 同一部分也提到:🛡 llama.cpp single = max robustness. Full 200K context on one 3090 (max-safe , fills cleanly with margin; see CLIFFS , slower than vLLM dual but doesn't crash on real-world tool-using agents.。