Your-First-LLM-Studio

First LLM Studio: local-first LLM studio for Apple Silicon with MLX runtimes, Compare Lab, benchmark ops, replay, and runtime telemetry.

101
Stars
8
Forks
TypeScript
Language
8/24/2026
Added
View on GitHubDownload ZIP

⚠️ Third-Party Software Notice

This skill is third-party open-source software developed and hosted independently on GitHub. SkillTip is an informational directory and does not control or maintain the underlying repository. Any security checks displayed are automated and limited in scope. Review the source code before installing.

Read the Terms of Service

Installation

Add to your Claude Code skills directory:

# Add to your Claude Code skills
git clone https://github.com/ChrisChen667788/Your-First-LLM-Studio

Getting Started

Guides for using skills like Your-First-LLM-Studio.

Security Report

Verified

Last scanned: —

{
  "status": "PASSED",
  "issues": []
}

README.md

First LLM Studio

English | 简体中文

Release License Apple Silicon MLX

First LLM Studio hero


English

First LLM Studio is a local-first LLM workbench for Apple Silicon. It brings local MLX runtimes, remote API targets, Agent sessions, Compare, Fine-tune, Benchmark, model discovery, runtime recovery, release evidence, and admin monitoring into one operating surface.

It is not another chat shell. It is built for people who need to compare behavior, debug runtimes, run evals, prepare adapters, and keep local and remote model work inside one product loop.

Product Surfaces

RouteCore workflow
/agentTool-enabled Agent sessions, target selection, runtime state, replay, trace review, and embedded Compare entry.
/compareRoute-owned Compare Studio for prompt composition, lane preview, recipe persistence, review drawer, and benchmark handoff.
/fine-tuneForeground Fine-tune Studio for datasets, recipes, training, evaluation, chat adapter proof loops, export, reports, and artifacts.
/modelsModel discovery and install verification for local/community models plus hardware-fit and risk signals.
/benchmarksBenchmark run controls, progress, reports, release evidence, baselines, and regression review.
/retrievalForeground knowledge management, path import, chunk inspection, and grounded retrieval validation.
/experimentsUnified run/session timeline with artifact lineage, cross-feature navigation, filters, and retention controls.
/adminMonitoring/configuration mirror for runtime, queues, benchmark history, provider health, guardrails, and audit timelines.

Major Version Story

VersionCore capabilities
v0.1 FoundationEstablished the local-first web studio, Apple Silicon/MLX gateway workflow, local + remote target catalog, runtime telemetry, and the first Agent/Admin operating split.
v0.2 Agent + Benchmark OpsAdded richer Agent workbench flows, Compare-style target review, replay/trace inspection, runtime recovery controls, formal benchmark operations, baselines, and regression evidence.
v0.3 Fine-tune + Release EvidenceAdded fine-tune operation loops for evaluation, adapter chat, adapter export, and distillation starters; expanded operation history, partitioned typechecks, screenshot smoke, route smoke, and public launch assets.
v0.4 Product IA releaseMoves /fine-tune, /compare, /models, /benchmarks, /retrieval, and /experiments into foreground product routes with feature-owned state/actions, artifact lineage and retention, dark-glass studio/workbench styling, canonical APIs, and admin narrowed toward monitoring/configuration.
v0.4.1 Stability baselineRepairs dataless workspace failure modes, keeps route smoke and typecheck green, updates the OpenAI-compatible /v1 surface, refreshes provider status reporting, captures current real UI evidence, and archives a real Qwen3 4B LoRA run with checkpoint/report/chart evidence.
v0.4.2 Evidence patchFormalizes the GitHub/ModelScope high-resolution screenshot sync, documents the README-facing LFS threshold fix, and preserves the v0.4.1 LoRA evidence as the stable public baseline while v0.5.0 work starts.
v0.5 Starter trackEnterprise RAG, deployment registry, OpenAI-compatible API, telemetry, release-readiness gates, production attestation, and control-plane rehearsal work continue behind explicit preview gates until promoted.
v1.0 Integrated GA baselineUnifies Agent, Compare, Model Hub, Retrieval, Fine-tune, Benchmark, Experiments, Admin monitoring, thin application APIs, route ownership, release security, and reproducible evidence contracts.
v1.1.0-rc.1 Desktop OnboardingAdds a self-contained Apple Silicon app with bundled Node, ZIP/DMG packaging, first-run diagnosis, permission and service recovery, migration/update/rollback/uninstall rehearsals, a real Ollama local-chat proof, and clean-profile boot evidence. Developer ID notarization remains a separate GA gate.
v1.1.0-rc.2 Desktop Distribution GateReplaces the shell entrypoint with a native arm64 launcher and adds nested-code/app/DMG signing, dual notarization logs, staple/Gatekeeper verification, a portable clean-machine runner, and RSA-signed organization receipts with an out-of-band trust pin. Real external receipts still gate GA.
v1.1.1 Model Hub LifecycleAdds immutable multi-file Hub manifests, provider SHA-256 receipts, operator-approved physical external-volume migration, ownership manifests, and a visible promotion read model. Refreshed Hub identity evidence remains a separate gate.
v1.2.0 Local Server AcceptanceAdds a real Ollama 15-slice acceptance loop for process health, model residency, OpenAI-compatible chat/SSE, concurrency, accounting, access policy, log retention, idle eviction, and unload/reload recovery. Separate-device LAN and sustained daemon evidence still gate production promotion.
v1.2.1 Runtime FabricImplements one normalized runtime contract across MLX, Ollama, llama.cpp, LocalAI, vLLM, and SGLang. Real MLX/Ollama/llama.cpp chat and SSE pass on Apple Silicon; unsupported hardware and missing endpoints fail before execution with actionable codes. External LocalAI, Linux/NVIDIA, and heterogeneous-node receipts still gate production promotion.
v1.3.0 MCP + Secure ExtensionsAdds a pinned MCP server registry, real stdio capability discovery, Ed25519-signed install/update/rollback, permission and secret scopes, quarantine, dependency/path defenses, and OS-enforced macOS Seatbelt isolation. Local acceptance is 11/11 PASS; independent publisher, Linux/Windows sandbox, and remote OAuth receipts still gate production promotion.
v1.3.1 Workflow Studio + Trust HardeningAdds typed visual workflows, immutable publication, safe-worker resume/replay, strict deployment-key invocation, standard OpenAI-compatible sync/SSE responses, atomic recoverable Workflow stores, conventional contract tests, zero-vulnerability dependency gates, and portable LoRA evidence with quality promotion kept on HOLD.
v1.4.0 Team GovernanceAdds organizations, workspaces, roles, groups, request identity, optimistic conflict handling, database-level ACL/RLS rehearsals, identity mapping, policy simulation, and immutable audit evidence. Real OIDC/SCIM and deployed PostgreSQL remain production gates.
v1.4.1 Quality and Training LabBinds frozen multi-seed evaluation, confidence intervals, deterministic scoring, Quality CI, and training capability contracts to a real attached adapter and 36 paired samples. Independent-worker repetition and calibrated subjective judging remain external gates.
v1.5.0 Trusted Artifact LifecyclePackages the real adapter with pinned base revision, checksums, provenance, local registry read-back, quality-claim binding, install policy, and rollback lifecycle. Organization-controlled remote registry receipts remain a production gate.
v1.5.1 Enterprise HA and FinOpsCompletes the local durable usage outbox, token reconciliation, retry-safe settlement, audit/signing rehearsal, old-primary fencing, standby promotion, and measured local RPO/RTO. Managed billing, cross-region failover, cloud KMS/Object Lock, and organization sign-off remain HOLD.
v1.6.3 Benchmark Qualification CheckpointPins and checksum-verifies the complete HuggingFaceH4/MATH-500 test snapshot, exposes all 500 qualified items in Benchmark Studio, and preserves reproducible dataset provenance.
v1.6.4 Official Evaluator CheckpointAdds pinned Math-Verify 0.9.0 equivalence scoring, resumable per-sample checkpoints, a real 500/500 local Qwen3 0.6B run at 32.00% accuracy, and protocol adapters for MMMU, MathVista, MMBench, and Video-MME v2. Full multimodal execution and external leaderboard reproduction remain HOLD.
v1.6.5 Benchmark Reproducibility CheckpointAdds subject/difficulty scorecards, Wilson 95% confidence, latency/token/failure accounting, immutable run and evaluator fingerprints, a 500/500 isolated scorer replay with 100% decision agreement, and executable multimodal readiness plans. The replay is same-host evidence; independent workers and official multimodal runs remain HOLD.
v1.6.6 Benchmark Decision IntelligenceConverts the real 500-item run into complete error taxonomy, confidence-aware cohort risks, latency/token outliers, a bounded review queue, conservative power planning, and a paired candidate gate with McNemar and non-inferiority policy. Local audit acceptance is separate from candidate promotion, which remains EVIDENCE NEEDED until a second distinct full run exists.

Current source version: v1.5.1, with source status COMPLETE and local acceptance PASS. The machine-readable release-state.json deliberately separates the active source milestone from the latest public GitHub release (v0.4.0) and the latest desktop candidate (v1.1.0-rc.2).

Internal post-release evidence has advanced through the v1.6.6 benchmark decision-intelligence checkpoint: all 500 stored results are classified (160 correct and 340 mathematical mismatches, with 100% answer extraction), confidence-aware risk highlights Intermediate Algebra, Precalculus, and Level 5, and the conservative 500-sample detectable effect is 8.27 percentage points. Only one distinct complete run exists, so candidate promotion remains EVIDENCE NEEDED. This does not change the active source version or claim hosted leaderboard parity, independent-host reproduction, multimodal execution, distribution, or production promotion.

Latest desktop candidate notes: v1.1.0-rc.2. Distribution remains HOLD and production remains BLOCKED; this repository does not represent missing Apple, organization, non-loopback, distributed-worker, or collaborative receipts as completed evidence.

Competitive Position

Reviewed against official product documentation on 2026-07-12. Core means a first-party primary workflow; integrated means the capability exists but is not the product's deepest specialization; ecosystem means it is normally assembled through adjacent clients or plugins. "Not primary" does not mean technically impossible.

ProductStrongest positionLocal runtime / Model HubAgent / RAGFine-tune / LoRAEvaluation / operational evidence
First LLM StudioEvidence-driven local model lifecycleCore, MLX and hardware-awareCore, tools + Compare + ACL/citationsCore, recipe-to-best-checkpoint-to-adapter lifecycleCore, Benchmark + lineage + fail-closed release gates
LM StudioPolished desktop discovery and local servingCore, GUI/CLI load, download, unload and compatible APIsIntegrated through API, tools and MCPNot primaryRuntime/developer inspection
OllamaSimple model runtime and packagingCore, stable local API and model lifecycleEcosystem, with native tool callingNot primaryEcosystem
Open WebUISelf-hosted team AI workspaceIntegrated, multi-providerCore, hybrid RAG, reranking, tools and MCPNot primaryIntegrated, arena/A-B/ELO, analytics and OTel
JanOpen-source cross-platform desktop assistantCore, llama.cpp/MLX and configurable local serverIntegrated, agents/projects/MCPNot primaryServer logs and developer inspection
AnythingLLMWorkspace RAG and agent automationIntegrated, multi-providerCore, workspaces, flows, skills and scheduled jobsNot primaryLogs and flow runs
LLaMA-FactoryEfficient training breadth and depthTraining/inference toolingTask-focused tool-use tuningCore, broad LoRA/QLoRA and preference methodsCore training monitors and benchmark integrations
LocalAIModular multi-backend private AI runtimeCore, broad hardware/backends and distributed workersCore, agents/MCP/RAG/citationsIntegratedRuntime and control-plane operations

First LLM Studio's advantage is the connected evidence chain across Agent, Compare, Retrieval, Benchmark, Fine-tune, adapter export, Workflow Studio, and release review. Its clearest remaining gaps are signed desktop distribution, heterogeneous runtime production receipts, team identity/governance, collaborative and distributed workflow execution, and real multi-node/cloud evidence. See the full bilingual competitive landscape and product direction, including methodology, limitations, official sources, and the post-v1 adoption principles.

Who It Helps

Local AI builders on Apple Silicon

  • Compare MLX local models against hosted APIs under aligned context budgets.
  • Inspect runtime cost, prewarm, release, recovery, and hardware pressure without leaving the app.
  • Decide which local model is actually usable for daily coding and analysis workflows.

Agent and tooling teams

  • Validate tool-calling, repo-grounded behavior, replay, and patch flows in one workbench.
  • Turn Compare runs into benchmark handoff without switching products.
  • Separate model-quality failures from provider quirks and local runtime instability.

Evaluation and platform engineers

  • Run formal and focused benchmark suites with repeatable profiles.
  • Review baselines, deltas, run notes, failure classifications, and release evidence.
  • Keep local and remote targets inside one comparable target catalog.

Core Value

  • Unified local + remote target catalog.
  • Compare Lab for model-vs-model output review.
  • Fine-tune workflows for datasets, recipes, training, evaluation, adapter proof loops, export, best-checkpoint selection, and LoRA release evidence.
  • Visual Workflow Studio for typed Agent/RAG/eval graphs, immutable recipes, guarded tool execution, replay, and OpenAI-compatible deployment.
  • Benchmark operations with history, progress, baselines, reports, and release evidence.
  • Foreground Retrieval for document import, chunk inspection, and grounded evidence probes.
  • Experiments timeline for session/run lineage, artifact navigation, and retention policy.
  • Replay, trace review, patch inspection, and exportable review notes.
  • Runtime operations for prewarm, release, restart, log inspection, telemetry, and recovery.
  • Dynamic local/community model discovery plus remote provider health scanning.

Current Targets

Local

  • Local Qwen3 0.6B
  • Local Qwen3 4B 4-bit
  • Local Qwen3.5 4B 4-bit
  • Local Gemma 3 4B It Qat 4-bit

Remote

  • OpenAI Codex
  • OpenAI GPT-5.5
  • Claude API
  • DeepSeek API
  • Kimi API
  • GLM API
  • Qwen API

Selection and stability guidance: docs/benchmark-lane-comparison.md.

Contributor onboarding: English · 中文快速上手 · GitHub setup checklist.

Screenshots

Captured from the running local app after npm run typecheck and npm run smoke:routes. README screenshots are generated with npm run screenshots:readme at 2x DPR so text stays sharp on GitHub and ModelScope; the LoRA evidence chart is exported from SVG at 2x DPR.

Agent workbench with target catalog, runtime rail, and tool-enabled composer:

Agent workbench

Workflow Graph Studio with draggable typed nodes, version/revision controls, execution recovery, and promotion evidence:

Workflow Graph Studio

Reproducible motion capture workflow: docs/demo-video-workflow.md.

Watch the Agent workbench MP4 demo · SHA-256 metadata

Fine-tune Studio with workflow tabs, training controls, and report/evidence panels:

Fine-tune Studio

Fine-tune completed run with live loss curves, train/validation traces, and handoff actions:

Fine-tune training curve

Real Qwen3 4B LoRA release evidence with save/eval markers and selected best checkpoint:

Qwen3 4B LoRA release evidence

Vector version: fine-tune-qwen4b-lora-chart.svg. Full run archive and manifest: docs/release-evidence/finetune-qwen4b-lora-2026-07-01.

Benchmark Studio with run controls and historical evidence cards:

Benchmark Studio

Benchmark run evidence from a real local smoke run:

Benchmark run evidence

Models Studio with immutable Hub/storage receipts, real Ollama Local Server acceptance, and the real MLX/Ollama/llama.cpp Runtime Fabric matrix:

Models Studio

MCP and secure extension acceptance with signed lifecycle, real tool discovery, quarantine defenses, and explicit production gates:

MCP and secure extension acceptance

Compare, Retrieval, and Admin surfaces:

Compare Studio Retrieval Studio Admin dashboard Admin benchmark heatmap

Quick Start

Requirements

  • macOS on Apple Silicon
  • Node 22.x
  • Python 3.12
  • MLX-compatible local environment

Install

nvm install 22
nvm use 22
npm install
cp .env.example .env.local

Start the web app

npm run dev

Default routes:

Start the local model gateway

python3 -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install mlx mlx-lm
python scripts/local_model_gateway_supervisor.py

If your preferred Python is outside PATH, set LOCAL_AGENT_PYTHON_BIN before starting the app or gateway.

Gateway health:

Verification

npm run typecheck:changed
npm run smoke:routes
npm run smoke:screenshots

Configuration

Copy .env.example to .env.local and fill only the providers you want to use.

Important notes:

  • .env.local is ignored by git.
  • Remote providers are optional.
  • Several targets use OpenAI-compatible or Claude-compatible endpoints.
  • Public defaults in this repository are sanitized placeholders.

Repository Structure

app/                      Next.js app routes and thin API transports
components/               Shared UI and compatibility shells
features/                 Feature-owned routes, contracts, state, actions, and application ports
lib/agent/                Agent runtime, providers, benchmark, gateway helpers
lib/finetune/             Fine-tune store facade and split operation services
scripts/                  Local gateway, runtime, verification, and release scripts
docs/                     Architecture, release notes, launch notes, roadmap, and assets
modelscope/               ModelScope profile/readme metadata
public/                   Public assets and social cover art

Distribution

The ModelScope package script exports the committed Git tree so GitHub and ModelScope can stay file-identical for each synced version.

Security and Privacy

  • Sensitive local actions require confirmation.
  • Secrets belong in .env.local.
  • Public repository defaults are sanitized.
  • New public commits should use a GitHub noreply address where possible.
  • See SECURITY.md.

Contributing

Issues and PRs are welcome.

Release Notes


简体中文

First LLM Studio 是一个面向 Apple Silicon 的本地优先 LLM 工作台。它把本地 MLX 运行时、远端 API 目标、Agent 会话、Compare 对比、Fine-tune 微调、Benchmark 评测、模型发现、runtime 恢复、发布证据和后台监控统一到一个产品界面里。

它不是另一个聊天壳,而是给真正需要比较模型行为、调试 runtime、跑评测、准备 adapter,并把本地/远端模型工作流收在同一个产品循环里的开发者使用。

产品入口

路由核心工作流
/agent带工具循环的 Agent 会话、target 选择、runtime 状态、replay、trace review,以及内嵌 Compare 入口。
/compare前台 Compare Studio,负责 prompt 编排、lane preview、recipe 持久化、review drawer 和 benchmark handoff。
/fine-tune前台 Fine-tune Studio,覆盖数据集、配方、训练、评估、adapter proof loop、导出、报告和 artifacts。
/models本地/社区模型发现、安装验证、硬件适配和风险提示。
/benchmarksBenchmark run controls、进度、报告、发布证据、baseline 和回归审阅。
/retrieval前台知识管理、路径导入、chunk 检查和 grounded retrieval 验证。
/experiments统一 Session/Run 时间线、artifact lineage、跨功能导航、筛选和保留策略。
/adminRuntime、队列、benchmark 历史、provider health、guardrails 和 audit timeline 的监控/配置镜像。

大版本核心功能

版本核心功能
v0.1 基础版建立本地优先 Web Studio、Apple Silicon/MLX 网关工作流、本地 + 远端 target catalog、runtime telemetry,以及 Agent/Admin 的第一版操作分层。
v0.2 Agent + Benchmark 运维增强 Agent 工作台、Compare 式 target review、replay/trace 检查、runtime recovery controls、正式 benchmark 运维、baseline 和回归证据。
v0.3 Fine-tune + 发布证据加入 evaluation、adapter chat、adapter export、distillation starter 等 fine-tune 操作循环;扩展 operation history、分区 typecheck、截图 smoke、route smoke 和公开发布素材。
v0.4 产品结构发布版/fine-tune/compare/models/benchmarks/retrieval/experiments 推进为前台产品路由;迁移 feature-owned state/actions;接通 artifact lineage、导航和 retention;统一 dark-glass studio/workbench 视觉;使用 canonical API;Admin 收口为监控/配置。
v0.4.1 稳定基线修复 dataless 工作区导致的启动/编译卡死,保持 route smoke 和 typecheck 通过,刷新 OpenAI-compatible /v1 接口、provider 状态回报和当前实机 UI 证据,并归档真实 Qwen3 4B LoRA 训练的 checkpoint/report/chart evidence。
v0.4.2 证据补丁正式固化 GitHub/ModelScope 高清截图同步、README 截图 LFS 阈值修复,并保留 v0.4.1 LoRA 证据作为稳定公开基线,同时启动 v0.5.0 开发。
v0.5 Starter 轨道企业 RAG、部署 registry、OpenAI-compatible API、telemetry、release-readiness gates、生产签收和 control-plane rehearsal 持续放在显式 preview gate 后推进,满足证据门槛后再 promotion。
v1.0 一体化 GA 基线统一 Agent、Compare、Model Hub、Retrieval、Fine-tune、Benchmark、Experiments、Admin 监控、thin application API、route ownership、release security 与可复现证据契约。
v1.1.0-rc.1 桌面首次启动加入自包含 Apple Silicon app、内置 Node、ZIP/DMG、首次诊断、权限与服务恢复、迁移/更新/回滚/卸载演练、真实 Ollama 本地对话证明和 clean-profile 启动证据;Developer ID notarization 继续作为独立 GA 门禁。
v1.1.0-rc.2 桌面分发门禁将 shell 入口替换为原生 arm64 launcher,并加入内部代码/app/DMG 分层签名、双层公证日志、staple/Gatekeeper 验证、独立 Mac 验收脚本及带线下信任锚的 RSA 组织签收;真实外部 receipt 仍是 GA 门禁。
v1.1.1 Model Hub 生命周期加入不可变多文件 Hub manifest、provider SHA-256 receipt、operator-approved 物理外置盘迁移、ownership manifest 和可视 promotion read model;更新后的 Hub identity receipt 仍是独立门禁。
v1.2.0 Local Server 验收加入真实 Ollama 15-slice 验收,覆盖进程健康、模型驻留、OpenAI-compatible chat/SSE、并发、计量、访问策略、日志保留、idle eviction 和 unload/reload recovery;跨设备 LAN 与持续 daemon 证据继续作为生产门禁。
v1.2.1 Runtime Fabric用同一标准化合同实现 MLX、Ollama、llama.cpp、LocalAI、vLLM 与 SGLang 适配器;Apple Silicon 上真实 MLX/Ollama/llama.cpp chat 与 SSE 全部通过,硬件或端点不满足时会在执行前给出可操作错误码;外部 LocalAI、Linux/NVIDIA 与异构节点 receipt 继续作为生产门禁。
v1.3.0 MCP + 安全扩展加入固定版本 MCP server registry、真实 stdio capability discovery、Ed25519 签名安装/升级/回滚、权限与密钥 scope、quarantine、依赖/路径防御及 macOS Seatbelt 强制隔离;本地验收 11/11 PASS,独立 publisher、Linux/Windows sandbox 与远程 OAuth receipt 继续作为生产门禁。
v1.3.1 Workflow Studio + 信任加固加入类型化可视工作流、不可变发布、安全 worker 恢复/回放、严格 deployment key、标准 OpenAI-compatible 同步/SSE、可原子恢复的 Workflow 存储、常规合同测试、零漏洞依赖门禁,以及保持质量晋级 HOLD 的可移植 LoRA 证据。
v1.4.0 团队治理加入组织、工作区、角色、用户组、请求身份、乐观并发冲突、数据库级 ACL/RLS 演练、身份映射、策略模拟与不可变审计证据;真实 OIDC/SCIM 与部署后的 PostgreSQL 仍是生产门禁。
v1.4.1 质量与训练实验室将冻结多种子评测、置信区间、确定性评分、Quality CI 和训练能力合同绑定到真实已挂载 adapter 与 36 组配对样本;独立 worker 重跑和主观 judge 校准仍是外部门禁。
v1.5.0 可信 Artifact 生命周期使用固定 base revision、checksum、provenance、本地 registry 回读、质量声明绑定、安装策略和回滚生命周期封装真实 adapter;组织控制的远端 registry receipt 仍是生产门禁。
v1.5.1 企业 HA 与 FinOps完成本地 durable usage outbox、token 对账、可安全重试的结算、审计/签名演练、旧主 fencing、standby promotion 与本地 RPO/RTO 测量;托管 billing、跨区 failover、云 KMS/Object Lock 和组织签收继续保持 HOLD
v1.6.3 Benchmark 资格化检查点固定并校验完整 HuggingFaceH4/MATH-500 test 快照,在 Benchmark Studio 暴露全部 500 个资格化样本,并保留可复现数据 provenance。
v1.6.4 官方判分器检查点接入固定版本 Math-Verify 0.9.0 等价判分、逐题可恢复 checkpoint,完成真实 Qwen3 0.6B 500/500 本地运行(32.00%),并提供 MMMU、MathVista、MMBench、Video-MME v2 协议 adapter;多模态全量执行与外部榜单复现继续保持 HOLD
v1.6.5 Benchmark 复现检查点增加学科/难度 scorecard、Wilson 95% 区间、延迟/token/失败分类、不可变 run 与 evaluator 指纹,并通过新隔离 scorer worker 对 500/500 输出重判且 decision 100% 一致;该结果仍是同机证据,独立 worker 与官方多模态全量运行继续 HOLD
v1.6.6 Benchmark 决策智能将真实 500 题 run 转成完整错误分类、置信区间 cohort 风险、延迟/token 异常点、有限复核队列、保守统计功效规划,以及带 McNemar 与非劣效策略的配对候选门槛。本地审计验收与候选晋级分开;在出现第二个不同 run id 的完整运行前,候选晋级保持 EVIDENCE NEEDED

当前源码版本:v1.5.1,源码状态为 COMPLETE、本地验收为 PASS。机器可读的 release-state.json 明确区分当前源码里程碑、最近公开 GitHub Release(v0.4.0)与最近桌面候选包(v1.1.0-rc.2)。

内部后续证据已推进到 v1.6.6 Benchmark 决策智能检查点:500 条结果全部完成分类(160 正确、340 条数学不等价,答案提取覆盖 100%),置信区间规则标出 Intermediate Algebra、Precalculus 与 Level 5 风险,500 样本下的保守可检测变化约为 8.27 个百分点。当前只有一个不同 run id 的完整运行,因此候选晋级继续为 EVIDENCE NEEDED。该检查点不改变当前源码版本,也不宣称托管榜单等价、独立主机复现、多模态全量执行、分发或生产晋级。

最新桌面候选说明:v1.1.0-rc.2。分发状态继续为 HOLD、生产状态继续为 BLOCKED;仓库不会把缺失的 Apple、组织、非回环、分布式 worker 或多人协作 receipt 表述为已完成证据。

竞品定位对比

本表基于 2026-07-12 可查的官方产品文档。核心表示产品原生主流程,已集成表示具备能力但不是最深的专长,生态表示通常依赖相邻客户端或插件组装。“非主线”不等于技术上完全不能实现。

产品最强定位本地运行时 / Model HubAgent / RAGFine-tune / LoRA评测 / 运维证据
First LLM Studio证据驱动的本地模型全生命周期核心,MLX 与硬件感知核心,工具 + Compare + ACL/引用核心,从 recipe、最佳 checkpoint 到 adapter lifecycle核心,Benchmark、lineage 与 fail-closed 发布门槛
LM Studio成熟的桌面模型发现和本地服务核心,GUI/CLI 加载、下载、卸载和兼容 API已集成,API、工具和 MCP非主线Runtime / Developer 检查
Ollama简洁稳定的模型运行时与打包核心,本地 API 与模型生命周期生态,原生支持 tool calling非主线生态提供
Open WebUI面向团队的自托管 AI 工作区已集成,多 provider核心,混合 RAG、reranker、工具和 MCP非主线已集成,Arena/A-B/ELO、分析和 OTel
Jan开源跨平台桌面助手核心,llama.cpp/MLX 与可配置 Local Server已集成,Agent/Project/MCP非主线Server 日志与开发检查
AnythingLLMWorkspace RAG 与 Agent 自动化已集成,多 provider核心,Workspace、Flow、Skill 和定时任务非主线日志与 Flow Run
LLaMA-Factory高效训练的模型/方法覆盖深度训练与推理工具面向任务的工具调用训练核心,广泛 LoRA/QLoRA 与偏好训练核心,训练监控与 benchmark 集成
LocalAI模块化、多后端的私有 AI runtime核心,广硬件/后端和分布式 worker核心,Agent/MCP/RAG/引用已集成Runtime 与控制面运维

First LLM Studio 的优势是把 Agent、Compare、Retrieval、Benchmark、Fine-tune、adapter export、Workflow Studio 和 release review 接成一条证据链。当前最明确的剩余短板是签名桌面分发、异构 runtime 生产证据、团队身份与治理、协同/分布式 workflow 执行,以及真实多节点/云生产证据。完整方法、优劣势、官方来源和 post-v1 借鉴原则见双语文档:竞品格局与产品方向

对哪些用户有价值

Apple Silicon 本地 AI 开发者

  • 在统一上下文预算下,对比 MLX 本地模型和托管 API。
  • 不离开应用就能查看 runtime 成本、prewarm、release、恢复动作和硬件压力。
  • 判断哪个本地模型真的适合日常 coding / analysis 工作流。

Agent / 工具链团队

  • 在一个工作台里验证 tool calling、repo-grounded behavior、replay 和 patch 流程。
  • 直接把 Compare 结果送入 Benchmark,不必切换产品。
  • 区分失败来源:模型质量、provider 行为,还是本地 runtime 不稳。

评测 / 平台工程团队

  • 用可复现 profile 跑 formal 和 focused benchmark suites。
  • 查看 baseline、delta、run note、失败分类和发布证据。
  • 让本地与远端 target 落在同一个可比较的 target catalog 里。

核心价值

  • 本地 + 远端统一 target catalog。
  • Compare Lab 支持模型对模型审阅。
  • Fine-tune 工作流覆盖 dataset、recipe、training、evaluation、adapter proof loop、export、best-checkpoint 选择和 LoRA 发布证据。
  • 可视化 Workflow Studio 覆盖 Agent/RAG/eval 类型图、不可变 recipe、受保护工具执行、回放与 OpenAI-compatible 部署。
  • Benchmark 运维覆盖 history、progress、baseline、report 和 release evidence。
  • Retrieval 前台覆盖文档导入、chunk 检查和 grounded evidence probe。
  • Experiments 时间线覆盖 Session/Run lineage、artifact 导航和 retention policy。
  • Replay、trace review、patch inspection 与可导出的审阅记录。
  • Runtime 运维覆盖 prewarm、release、restart、日志检查、telemetry 和 recovery。
  • 支持本地/社区模型发现和远端 provider health 扫描。

当前支持的 Target

本地

  • Local Qwen3 0.6B
  • Local Qwen3 4B 4-bit
  • Local Qwen3.5 4B 4-bit
  • Local Gemma 3 4B It Qat 4-bit

远端

  • OpenAI Codex
  • OpenAI GPT-5.5
  • Claude API
  • DeepSeek API
  • Kimi API
  • GLM API
  • Qwen API

Target 选择、稳定性与适用任务对照:docs/benchmark-lane-comparison.md

贡献者入口:English · 中文快速上手 · GitHub 仓库设置清单

截图

以下截图来自本地运行版本,并已通过 npm run typechecknpm run smoke:routes。README 截图使用 npm run screenshots:readme 以 2x DPR 生成,LoRA 证据图从 SVG 以 2x DPR 导出,确保在 GitHub 和 ModelScope 缩放后文字仍然清晰。

Agent 工作台:target catalog、runtime rail 与工具化输入区:

Agent 工作台

Workflow Graph Studio:可拖拽类型节点、版本/修订控制、执行恢复与 promotion evidence:

Workflow Graph Studio

可复现动态演示流程:docs/demo-video-workflow.md

查看 Agent 工作台 MP4 演示 · SHA-256 元数据

Fine-tune Studio:工作流 tab、训练控制与 report/evidence 面板:

Fine-tune Studio

Fine-tune 完成作业:真实 loss 曲线、训练/验证轨迹与 handoff 操作:

Fine-tune 训练曲线

真实 Qwen3 4B LoRA 发布证据:包含 save/eval 事件标记和自动选择的最佳 checkpoint:

Qwen3 4B LoRA 发布证据

矢量版本:fine-tune-qwen4b-lora-chart.svg。完整 run archive 与 manifest:docs/release-evidence/finetune-qwen4b-lora-2026-07-01

Benchmark Studio:运行控制与历史证据卡片:

Benchmark Studio

Benchmark:本地 smoke run 生成的真实评测证据:

Benchmark 运行证据

Models Studio:不可变 Hub/外置盘证据、真实 Ollama Local Server 验收,以及真实 MLX/Ollama/llama.cpp Runtime Fabric 矩阵:

Models Studio

MCP 与安全扩展验收:签名生命周期、真实工具发现、隔离/检疫防御和显式生产门禁:

MCP 与安全扩展验收

Compare、Retrieval 与 Admin:

Compare Studio Retrieval Studio Admin dashboard Admin benchmark 热力图

快速开始

环境要求

  • Apple Silicon macOS
  • Node 22.x
  • Python 3.12
  • 可运行 MLX 的本地环境

安装

nvm install 22
nvm use 22
npm install
cp .env.example .env.local

启动 Web 应用

npm run dev

默认入口:

启动本地模型网关

python3 -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install mlx mlx-lm
python scripts/local_model_gateway_supervisor.py

如果 Python 不在 PATH 中,可以在启动前设置 LOCAL_AGENT_PYTHON_BIN 指向实际解释器。

网关健康检查:

验证

npm run typecheck:changed
npm run smoke:routes
npm run smoke:screenshots

配置说明

.env.example 复制成 .env.local,只填写你要启用的 provider 即可。

需要注意:

  • .env.local 已被 git 忽略。
  • 远端 provider 是可选的。
  • 部分 target 走 OpenAI-compatible / Claude-compatible endpoint。
  • 本仓库公开版本已经做过脱敏,占位值需要替换成你自己的 endpoint。

仓库结构

app/                      Next.js app routes 与 thin API transports
components/               共享 UI 与兼容 shell
features/                 feature-owned routes、contracts、state、actions、application ports
lib/agent/                Agent runtime、providers、benchmark、gateway helpers
lib/finetune/             Fine-tune store facade 与拆分后的 operation services
scripts/                  本地网关、runtime、验证和发布脚本
docs/                     架构、release notes、launch notes、roadmap 和 assets
modelscope/               ModelScope 主页/readme 元数据
public/                   对外资源和社媒封面图

发布与同步

ModelScope 打包脚本会导出已提交的 Git tree,因此每次同步都可以让 GitHub 和 ModelScope 保持同一份文件快照。

安全和隐私

  • 敏感本地操作默认需要确认。
  • Secret 应保存在 .env.local
  • 公开仓库默认配置已经做过脱敏。
  • 新增公开提交应尽量使用 GitHub noreply 地址。
  • SECURITY.md

贡献

欢迎 issue 和 PR。

发布说明

Frequently Asked Questions

What is Your-First-LLM-Studio?

Your-First-LLM-Studio is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by ChrisChen667788. First LLM Studio: local-first LLM studio for Apple Silicon with MLX runtimes, Compare Lab, benchmark ops, replay, and runtime telemetry. It has 101 GitHub stars.

Is Your-First-LLM-Studio safe to use?

Your-First-LLM-Studio returned warnings in SkillsLLM's automated security scan. It has no critical vulnerabilities, but review the flagged issues in the Security Report section before adding it to your workflow.

How do I install Your-First-LLM-Studio?

Clone the repository with "git clone https://github.com/ChrisChen667788/Your-First-LLM-Studio" and add it to your Claude Code skills directory (see the Installation section above).

What programming language is Your-First-LLM-Studio written in?

Your-First-LLM-Studio is primarily written in TypeScript. It is open-source under ChrisChen667788 on GitHub, so you can review or fork the full source.

Are there alternatives to Your-First-LLM-Studio?

Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh Your-First-LLM-Studio against similar tools.

Comments (0)

No comments yet. Be the first to share your thoughts!

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI Agentsai-agentsanthropicclaude-code
View details
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI Agentsai-agentsbrainstorming
View details

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI Agentsai-agentsanthropicclaude-code
View details

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI Agentsclaude-codeai-tools
View details

claude-code

by anthropics

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

120,03119,897Shell
AI Agents
View details

Developers Also Liked

Based on votes and bookmarks from developers who liked this skill

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI Agentsai-agentsanthropicclaude-code
View details
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI Agentsai-agentsbrainstorming
View details

n8n

by n8n-io

12

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

201,88160,308TypeScript
MCP Serversapisai-tools
View details

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI Agentsai-agentsanthropicclaude-code
View details

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI Agentsclaude-codeai-tools
View details