README.md

Meeting Intelligence

Pipeline

Local-first meeting intelligence — turns audio, video, and media URLs into timestamped transcripts, translations, and grounded meeting protocols with decisions, assignments, and risks.

Runs locally by default. Cloud LLMs require explicit opt-in.

Table of contents


Repository layout

meeting-intelligence/
├── plugin.json / plugin.yaml   — Hermes plugin manifest
├── __init__.py                 — register(ctx) shim (src-layout bridge)
├── src/meeting_intelligence/   — код: pipeline, protocol/, llm_endpoint,
│                                 telegram*.py (омниканальность), web/, mcp_server
├── dashboard/plugin_api.py     — REST-роутер дашборда
├── com.hermes.desktop/         — снапшот виджета (источник — репо Штурмана)
├── skills/                     — SKILL.md для агента (meeting-intelligence, dossier)
├── tests/                      — pytest (197 passed / 1 pre-existing)
├── adr/ + docs/                — ADR-001…016, schema.md, план фазы
├── models/pyannote/            — офлайн-бандл диаризации (~32 МБ)
├── scripts/                    — деплой, обслуживание
└── references/                 — образцы/референсы для генерации документов

Submodule в репо Штурмана: подключается как shturman/meeting-intelligence, после чистой установки подтягивать git -C <plugins>/shturman submodule update --init --depth 1.

Requirements

  • Python 3.10+
  • ffmpeg on PATH
  • ≥ 8 GB RAM
  • An LLM backend: LM Studio, Ollama, llama.cpp, or a cloud API (opt-in)
  • Optional: NVIDIA CUDA GPU for faster transcription
Platform Install ffmpeg
Windows winget install Gyan.FFmpeg
Linux sudo apt update && sudo apt install ffmpeg
macOS brew install ffmpeg

Install

Meeting Intelligence can be used three ways: as a standalone CLI, as a Hermes native plugin, or as a portable Agent Plugin (MCP).

Option A — Standalone CLI

python -m venv .venv && source .venv/bin/activate   # Windows: .\.venv\Scripts\Activate.ps1
# The package is NOT on PyPI — install from the repository:
pip install "meeting-intelligence[local] @ git+ssh://git@gitflic.ru/manve-sulimo2/shturman-ai.git"

Optional extras:

Extra What it adds
local Local LLM backend support (core deps already cover this).
cloud Explicit cloud LLM intent.
gpu NVIDIA CUDA runtime (Windows, Linux).
diarization Speaker diarization via pyannote.audio.
url Media URL download via yt-dlp.
web Web dashboard server (fastapi, uvicorn, python-multipart).
mcp MCP stdio server for Agent Plugins (mcp SDK).
all Everything above combined.
dev Test, lint, and build tools.

Option B — Hermes plugin

Install directly from Git — Hermes clones the repo, registers the plugin, and adds its toolset to the agent:

hermes plugins install git+ssh://git@gitflic.ru/manve-sulimo2/shturman-ai.git --enable
hermes gateway restart

Verify:

hermes plugins list          # meeting-intelligence … enabled … 0.8.0 … git
hermes tools list            # ✓ enabled  meeting_intelligence

Update to the latest version:

hermes plugins update meeting-intelligence
hermes gateway restart

Remove:

hermes plugins remove meeting-intelligence

Option C — Portable Agent Plugin (MCP)

For any agent-plugins.org v1.0.0 compatible client (Claude Code, Cursor, etc.):

  1. Clone or download this repo.
  2. The client discovers plugin.json, loads skills/meeting-intelligence/SKILL.md, and starts the MCP server declared in mcp.json:
{
  "$schema": "https://agent-plugins.org/schemas/1.0.0/mcp.schema.json",
  "mcpServers": {
    "meeting-intelligence": {
      "type": "stdio",
      "command": "python",
      "args": ["-m", "meeting_intelligence.mcp_server"]
    }
  }
}

The MCP server exposes eight tools: meeting_transcribe, meeting_translate, meeting_agent_transcript, meeting_protocol, meeting_process, meeting_enroll, meeting_voiceprints, meeting_dossier.

Heavy dependencies (faster-whisper, CUDA, ffmpeg) must be installed in the Python environment that runs the server.


LLM backend configuration (ADR-016)

There are no default models or URLs in the code. Every LLM call resolves at call time: MEETING_LLM_* env → shturman.yaml llm.model + llm.gateway.{base_url, api_key_env} → fail-closed with instructions.

Set three environment variables before translating or generating a protocol. LM Studio:

export MEETING_LLM_BASE_URL=http://localhost:1234/v1
export MEETING_LLM_API_KEY=lm-studio
export MEETING_LLM_MODEL=<your-model>

Ollama:

export MEETING_LLM_BASE_URL=http://localhost:11434/v1
export MEETING_LLM_API_KEY=ollama
export MEETING_LLM_MODEL=<your-model>

Cloud (opt-in): cloud endpoints are blocked unless the run explicitly passes --allow-cloud (CLI) / allow_cloud=True (API). The env does not grant cloud access by design:

export MEETING_LLM_BASE_URL=https://api.openai.com/v1
export MEETING_LLM_API_KEY="$OPENAI_API_KEY"
export MEETING_LLM_MODEL=gpt-4o-mini
# + pass --allow-cloud to the command

GLM models (z.ai) are handled automatically: reasoning thinking is disabled at the transport layer (reasoning tokens otherwise consume the entire max_tokens budget and truncate JSON output).

Vision / OCR backends (photo posts, ADR-012)

Photo posts and local images are described by a local vision model — no LM Studio required. Two standalone llama.cpp servers, launched by one script:

python scripts/start_vision_servers.py   # vision :8018 (LFM2.5-VL-3B) + ocr :8017 (OvisOCR2)
# vision via dedicated llama.cpp (recommended; survives LM Studio restarts)
export MEETING_VISION_BASE_URL=http://127.0.0.1:8018/v1
export MEETING_VISION_MODEL=lfm2.5-vl-3b

# document/scan mode instead of photo description:
export MEETING_VISION_MODE=ocr           # uses :8017 OvisOCR2 → Markdown

# fully without LM Studio: point the protocol LLM at llama.cpp too
# export MEETING_LLM_BASE_URL=http://127.0.0.1:8019/v1  (launch your text model there)

Without MEETING_VISION_BASE_URL the vl mode falls back to MEETING_LLM_BASE_URL (LM Studio default) — old behavior preserved.

PowerShell: replace export NAME=value with $env:NAME = "value".


Usage

CLI commands

# Transcribe audio/video → timestamped transcript
meeting transcribe /path/to/meeting.mp4 --model small --language en --device cpu

# Translate transcript → target language
meeting translate /path/to/meeting.transcript.txt --target-lang ru

# Extract protocol (decisions, assignments, risks) from transcript
meeting protocol /path/to/meeting.transcript.txt --model qwen2.5-7b-instruct --docx

# Full pipeline in one step: transcribe → translate → protocol
meeting process /path/to/meeting.mp4 \
  --stt-model small \
  --llm-model qwen2.5-7b-instruct \
  --language en \
  --target-lang ru \
  --docx

# Clean transcript for agent analysis (no LLM call)
meeting agent-transcript /path/to/meeting.transcript.txt

# Launch the web dashboard
meeting serve --host 127.0.0.1 --port 8000

SOURCE may be a local audio/video file or, with the url extra installed, a supported media URL (YouTube, direct links, etc.).

GPU acceleration:

pip install 'meeting-intelligence[gpu]'
meeting transcribe meeting.mp4 --device cuda

The CLI falls back to CPU if no usable GPU is found. macOS uses CPU only.

Web dashboard

Start the server (requires the web extra):

pip install 'meeting-intelligence[web]'
meeting serve                    # http://127.0.0.1:8000

Using the dashboard:

  1. Open http://127.0.0.1:8000 in your browser.
  2. Choose input mode:
    • Upload File — select a local audio/video file.
    • Paste URL — enter a media URL (requires url extra / yt-dlp).
  3. Select options: STT model, source language, translation target.
  4. Click Process Meeting.
  5. The pipeline runs in the background. Status updates automatically every 2 seconds (pending → running → done / error).
  6. When done, download the transcript, translation, protocol JSON, and DOCX.

The dashboard is designed for local single-user use. It binds to 127.0.0.1 by default and has no authentication.

API endpoints (for programmatic access):

Method Path Description
GET / Dashboard HTML page
POST /api/jobs Create job — multipart file upload or form field url
GET /api/jobs/{id} Poll job status and result
GET /api/jobs/{id}/files/{filename} Download an output file

Example — create a job from a URL:

curl -X POST http://127.0.0.1:8000/api/jobs \
  -F "url=https://example.com/meeting.mp4" \
  -F "stt_model=small" \
  -F "language=en" \
  -F "target_lang=ru"
# → 202 {"id": "a1b2c3d4e5f6", "status": "pending", ...}

curl http://127.0.0.1:8000/api/jobs/a1b2c3d4e5f6
# → {"id": "a1b2c3d4e5f6", "status": "done", "result": {"files": [...]}}

curl -OJ http://127.0.0.1:8000/api/jobs/a1b2c3d4e5f6/files/meeting.protocol.json

Inside a Hermes agent

Once installed via hermes plugins install, the agent has access to eight tools under the meeting_intelligence toolset. Just ask in natural language:

«Обработай запись встречи» (with an attached audio file)

«Сделай протокол по этой ссылке: https://…»

«Переведи транскрипт на русский»

The agent uses the meeting-intelligence skill (skills/meeting-intelligence/SKILL.md) to orchestrate the pipeline: it reads the transcript, enriches it with corporate context (Jira/Confluence/Email/Calendar via MCP), and produces a grounded protocol with source_quote validation.

Available tools:

Tool What it does
meeting_transcribe Audio/video → timestamped transcript
meeting_translate Translate transcript lines to target language
meeting_agent_transcript Clean transcript into agent-ready JSON (no LLM)
meeting_protocol Extract validated protocol (decisions, assignments, risks)
meeting_process Full pipeline: transcribe → translate → protocol

Pipeline API (Python)

For custom integrations, import the typed pipeline API directly:

from meeting_intelligence.pipeline import (
    transcribe, translate, protocol, process,
    TranscribeParams, ProtocolParams, ProcessParams,
)

# 1. Transcribe
result = transcribe(TranscribeParams(
    source="meeting.mp4",
    model="small",
    language="en",
    device="cpu",
))
print(result.transcript_path)     # Path to saved transcript
print(result.transcript[:200])    # First 200 chars
print(result.meta)                # {duration, language, segment_count, ...}

# 2. Extract protocol
proto = protocol(ProtocolParams(
    transcript=result.transcript_path,
    model="qwen2.5-7b-instruct",
    docx=True,
))
print(proto.valid)                # True/False
print(proto.protocol_path)        # Path to protocol.json
print(proto.validation)           # {errors, warnings, overall_confidence}

# 3. Full pipeline in one call
result = process(ProcessParams(
    source="meeting.mp4",
    stt_model="small",
    llm_model="qwen2.5-7b-instruct",
    language="en",
    target_lang="ru",
    docx=True,
))

Every function returns a typed dataclass (TranscribeResult, ProtocolResult, ProcessResult) — no argparse, no stdout capture.


Environment variables

Variable Default Description
MEETING_LLM_BASE_URL (none — fail-closed) LLM API endpoint (required for LLM stages)
MEETING_VISION_BASE_URL falls back to MEETING_LLM_BASE_URL Dedicated vision endpoint (standalone llama.cpp :8018)
MEETING_VISION_MODEL lfm2.5-vl-3b Vision model name for the vl backend
MEETING_VISION_MODE vl vl (photo description) or ocr (document → Markdown, OvisOCR2 :8017)
MEETING_OCR_BASE_URL http://127.0.0.1:8017/v1 OCR backend endpoint
MEETING_LLM_API_KEY (none) LLM API key
MEETING_LLM_MODEL (none — fail-closed) LLM model name
MEETING_VERIFY_BASE_URL / _API_KEY / _MODEL (none — uses main LLM) Tiered second-pass verification endpoint
MEETING_PROTOCOL_VERIFY true false disables second-pass protocol verification
MEETING_PROTOCOL_CHUNK_OVERLAP_LINES overlap of chunking Lines of overlap between protocol chunks
MEETING_ROOT platform default (~/Meeting win / ~/meeting posix) Meetings storage root
MEETING_TG_API_ID / MEETING_TG_API_HASH (none) Telegram API credentials for voice/video ingest
MEETING_TRANSCRIBE_MODEL auto (small / large-v3-turbo on CUDA) Default Whisper model
MEETING_TRANSCRIBE_DEVICE auto cpu or cuda
MEETING_TRANSCRIBE_COMPUTE int8 / float16 (CUDA) Whisper compute type
MEETING_TRANSCRIBE_LANG ru Default source language
MEETING_TRANSLATE_BATCH_SIZE 8 Lines per LLM translation batch
MEETING_MAX_FILE_MB 2048 Max input file size
MEETING_MAX_DURATION_SEC 7200 Max media duration
MEETING_PROTOCOL_CHUNK_SIZE 6000 Token threshold for protocol chunking
MEETING_AGENT_MODE false If true, CLI prints agent-ready JSON
MEETING_YT_PROXY socks5://127.0.0.1:12334 (YouTube only) Proxy for yt-dlp downloads; unset = direct

Safety

  • Cloud blocked by default. External endpoints are rejected unless --allow-cloud (CLI) / allow_cloud=True (API) is supplied. There is no env override by design — cloud access is a per-run explicit decision.
  • Grounded protocols. Every decision and assignment requires a source_quote verified against the transcript. Unverifiable items are flagged with warnings; fabricated items cause validation failure.
  • No secret logging. API keys and credentials are never written to logs.
  • Audit metadata. Each protocol includes source_hash, stt_model, llm_model, created_at, and cloud_allowed for traceability.

Development

pip install -e '.[dev]'
pytest -q                          # 202 tests (1 xfail host-specific)

The dev extra pulls everything the test suite imports (fastapi, mcp<2, openpyxl, telethon, PySocks, pytest). Verified on a clean Ubuntu 24.04 VM: without it the suite reports ~13 collection errors, with it — 201 passed / 1 skipped.

Architecture decisions: docs/adr/ · Design doc: docs/sdd-v0.8.0.md


License

MIT

Описание
Конвейеры
0 успешных
0 с ошибкой
Разработчики