Meeting Intelligence

Local-first meeting intelligence — turns audio, video, and media URLs into timestamped transcripts, translations, and grounded meeting protocols with decisions, assignments, and risks.
Runs locally by default. Cloud LLMs require explicit opt-in.
Table of contents
- Requirements
- Install
- LLM backend configuration
- Usage
- Environment variables
- Safety
- Development
- License
Repository layout
meeting-intelligence/
├── plugin.json / plugin.yaml — Hermes plugin manifest
├── __init__.py — register(ctx) shim (src-layout bridge)
├── src/meeting_intelligence/ — код: pipeline, protocol/, llm_endpoint,
│ telegram*.py (омниканальность), web/, mcp_server
├── dashboard/plugin_api.py — REST-роутер дашборда
├── com.hermes.desktop/ — снапшот виджета (источник — репо Штурмана)
├── skills/ — SKILL.md для агента (meeting-intelligence, dossier)
├── tests/ — pytest (197 passed / 1 pre-existing)
├── adr/ + docs/ — ADR-001…016, schema.md, план фазы
├── models/pyannote/ — офлайн-бандл диаризации (~32 МБ)
├── scripts/ — деплой, обслуживание
└── references/ — образцы/референсы для генерации документов
Submodule в репо Штурмана: подключается как
shturman/meeting-intelligence, после чистой установки подтягиватьgit -C <plugins>/shturman submodule update --init --depth 1.
Requirements
- Python 3.10+
ffmpegonPATH- ≥ 8 GB RAM
- An LLM backend: LM Studio, Ollama, llama.cpp, or a cloud API (opt-in)
- Optional: NVIDIA CUDA GPU for faster transcription
| Platform | Install ffmpeg |
|---|---|
| Windows | winget install Gyan.FFmpeg |
| Linux | sudo apt update && sudo apt install ffmpeg |
| macOS | brew install ffmpeg |
Install
Meeting Intelligence can be used three ways: as a standalone CLI, as a Hermes native plugin, or as a portable Agent Plugin (MCP).
Option A — Standalone CLI
python -m venv .venv && source .venv/bin/activate # Windows: .\.venv\Scripts\Activate.ps1
# The package is NOT on PyPI — install from the repository:
pip install "meeting-intelligence[local] @ git+ssh://git@gitflic.ru/manve-sulimo2/shturman-ai.git"
Optional extras:
| Extra | What it adds |
|---|---|
local |
Local LLM backend support (core deps already cover this). |
cloud |
Explicit cloud LLM intent. |
gpu |
NVIDIA CUDA runtime (Windows, Linux). |
diarization |
Speaker diarization via pyannote.audio. |
url |
Media URL download via yt-dlp. |
web |
Web dashboard server (fastapi, uvicorn, python-multipart). |
mcp |
MCP stdio server for Agent Plugins (mcp SDK). |
all |
Everything above combined. |
dev |
Test, lint, and build tools. |
Option B — Hermes plugin
Install directly from Git — Hermes clones the repo, registers the plugin, and adds its toolset to the agent:
hermes plugins install git+ssh://git@gitflic.ru/manve-sulimo2/shturman-ai.git --enable
hermes gateway restart
Verify:
hermes plugins list # meeting-intelligence … enabled … 0.8.0 … git
hermes tools list # ✓ enabled meeting_intelligence
Update to the latest version:
hermes plugins update meeting-intelligence
hermes gateway restart
Remove:
hermes plugins remove meeting-intelligence
Option C — Portable Agent Plugin (MCP)
For any agent-plugins.org v1.0.0 compatible client (Claude Code, Cursor, etc.):
- Clone or download this repo.
- The client discovers
plugin.json, loadsskills/meeting-intelligence/SKILL.md, and starts the MCP server declared inmcp.json:
{
"$schema": "https://agent-plugins.org/schemas/1.0.0/mcp.schema.json",
"mcpServers": {
"meeting-intelligence": {
"type": "stdio",
"command": "python",
"args": ["-m", "meeting_intelligence.mcp_server"]
}
}
}
The MCP server exposes eight tools: meeting_transcribe, meeting_translate, meeting_agent_transcript, meeting_protocol, meeting_process, meeting_enroll, meeting_voiceprints, meeting_dossier.
Heavy dependencies (faster-whisper, CUDA, ffmpeg) must be installed in the Python environment that runs the server.
LLM backend configuration (ADR-016)
There are no default models or URLs in the code. Every LLM call resolves at call time: MEETING_LLM_* env → shturman.yaml llm.model + llm.gateway.{base_url, api_key_env} → fail-closed with instructions.
Set three environment variables before translating or generating a protocol. LM Studio:
export MEETING_LLM_BASE_URL=http://localhost:1234/v1
export MEETING_LLM_API_KEY=lm-studio
export MEETING_LLM_MODEL=<your-model>
Ollama:
export MEETING_LLM_BASE_URL=http://localhost:11434/v1
export MEETING_LLM_API_KEY=ollama
export MEETING_LLM_MODEL=<your-model>
Cloud (opt-in): cloud endpoints are blocked unless the run explicitly passes --allow-cloud (CLI) / allow_cloud=True (API). The env does not grant cloud access by design:
export MEETING_LLM_BASE_URL=https://api.openai.com/v1
export MEETING_LLM_API_KEY="$OPENAI_API_KEY"
export MEETING_LLM_MODEL=gpt-4o-mini
# + pass --allow-cloud to the command
GLM models (z.ai) are handled automatically: reasoning thinking is disabled at the transport layer (reasoning tokens otherwise consume the entire max_tokens budget and truncate JSON output).
Vision / OCR backends (photo posts, ADR-012)
Photo posts and local images are described by a local vision model — no LM Studio required. Two standalone llama.cpp servers, launched by one script:
python scripts/start_vision_servers.py # vision :8018 (LFM2.5-VL-3B) + ocr :8017 (OvisOCR2)
# vision via dedicated llama.cpp (recommended; survives LM Studio restarts)
export MEETING_VISION_BASE_URL=http://127.0.0.1:8018/v1
export MEETING_VISION_MODEL=lfm2.5-vl-3b
# document/scan mode instead of photo description:
export MEETING_VISION_MODE=ocr # uses :8017 OvisOCR2 → Markdown
# fully without LM Studio: point the protocol LLM at llama.cpp too
# export MEETING_LLM_BASE_URL=http://127.0.0.1:8019/v1 (launch your text model there)
Without MEETING_VISION_BASE_URL the vl mode falls back to MEETING_LLM_BASE_URL (LM Studio default) — old behavior preserved.
PowerShell: replace
export NAME=valuewith$env:NAME = "value".
Usage
CLI commands
# Transcribe audio/video → timestamped transcript
meeting transcribe /path/to/meeting.mp4 --model small --language en --device cpu
# Translate transcript → target language
meeting translate /path/to/meeting.transcript.txt --target-lang ru
# Extract protocol (decisions, assignments, risks) from transcript
meeting protocol /path/to/meeting.transcript.txt --model qwen2.5-7b-instruct --docx
# Full pipeline in one step: transcribe → translate → protocol
meeting process /path/to/meeting.mp4 \
--stt-model small \
--llm-model qwen2.5-7b-instruct \
--language en \
--target-lang ru \
--docx
# Clean transcript for agent analysis (no LLM call)
meeting agent-transcript /path/to/meeting.transcript.txt
# Launch the web dashboard
meeting serve --host 127.0.0.1 --port 8000
SOURCE may be a local audio/video file or, with the url extra installed, a supported media URL (YouTube, direct links, etc.).
GPU acceleration:
pip install 'meeting-intelligence[gpu]'
meeting transcribe meeting.mp4 --device cuda
The CLI falls back to CPU if no usable GPU is found. macOS uses CPU only.
Web dashboard
Start the server (requires the web extra):
pip install 'meeting-intelligence[web]'
meeting serve # http://127.0.0.1:8000
Using the dashboard:
- Open
http://127.0.0.1:8000in your browser. - Choose input mode:
- Upload File — select a local audio/video file.
- Paste URL — enter a media URL (requires
urlextra / yt-dlp).
- Select options: STT model, source language, translation target.
- Click Process Meeting.
- The pipeline runs in the background. Status updates automatically every 2 seconds (pending → running → done / error).
- When done, download the transcript, translation, protocol JSON, and DOCX.
The dashboard is designed for local single-user use. It binds to 127.0.0.1 by default and has no authentication.
API endpoints (for programmatic access):
| Method | Path | Description |
|---|---|---|
GET |
/ |
Dashboard HTML page |
POST |
/api/jobs |
Create job — multipart file upload or form field url |
GET |
/api/jobs/{id} |
Poll job status and result |
GET |
/api/jobs/{id}/files/{filename} |
Download an output file |
Example — create a job from a URL:
curl -X POST http://127.0.0.1:8000/api/jobs \
-F "url=https://example.com/meeting.mp4" \
-F "stt_model=small" \
-F "language=en" \
-F "target_lang=ru"
# → 202 {"id": "a1b2c3d4e5f6", "status": "pending", ...}
curl http://127.0.0.1:8000/api/jobs/a1b2c3d4e5f6
# → {"id": "a1b2c3d4e5f6", "status": "done", "result": {"files": [...]}}
curl -OJ http://127.0.0.1:8000/api/jobs/a1b2c3d4e5f6/files/meeting.protocol.json
Inside a Hermes agent
Once installed via hermes plugins install, the agent has access to eight tools under the meeting_intelligence toolset. Just ask in natural language:
«Обработай запись встречи» (with an attached audio file)
«Сделай протокол по этой ссылке: https://…»
«Переведи транскрипт на русский»
The agent uses the meeting-intelligence skill (skills/meeting-intelligence/SKILL.md) to orchestrate the pipeline: it reads the transcript, enriches it with corporate context (Jira/Confluence/Email/Calendar via MCP), and produces a grounded protocol with source_quote validation.
Available tools:
| Tool | What it does |
|---|---|
meeting_transcribe |
Audio/video → timestamped transcript |
meeting_translate |
Translate transcript lines to target language |
meeting_agent_transcript |
Clean transcript into agent-ready JSON (no LLM) |
meeting_protocol |
Extract validated protocol (decisions, assignments, risks) |
meeting_process |
Full pipeline: transcribe → translate → protocol |
Pipeline API (Python)
For custom integrations, import the typed pipeline API directly:
from meeting_intelligence.pipeline import (
transcribe, translate, protocol, process,
TranscribeParams, ProtocolParams, ProcessParams,
)
# 1. Transcribe
result = transcribe(TranscribeParams(
source="meeting.mp4",
model="small",
language="en",
device="cpu",
))
print(result.transcript_path) # Path to saved transcript
print(result.transcript[:200]) # First 200 chars
print(result.meta) # {duration, language, segment_count, ...}
# 2. Extract protocol
proto = protocol(ProtocolParams(
transcript=result.transcript_path,
model="qwen2.5-7b-instruct",
docx=True,
))
print(proto.valid) # True/False
print(proto.protocol_path) # Path to protocol.json
print(proto.validation) # {errors, warnings, overall_confidence}
# 3. Full pipeline in one call
result = process(ProcessParams(
source="meeting.mp4",
stt_model="small",
llm_model="qwen2.5-7b-instruct",
language="en",
target_lang="ru",
docx=True,
))
Every function returns a typed dataclass (TranscribeResult, ProtocolResult, ProcessResult) — no argparse, no stdout capture.
Environment variables
| Variable | Default | Description |
|---|---|---|
MEETING_LLM_BASE_URL |
(none — fail-closed) | LLM API endpoint (required for LLM stages) |
MEETING_VISION_BASE_URL |
falls back to MEETING_LLM_BASE_URL |
Dedicated vision endpoint (standalone llama.cpp :8018) |
MEETING_VISION_MODEL |
lfm2.5-vl-3b |
Vision model name for the vl backend |
MEETING_VISION_MODE |
vl |
vl (photo description) or ocr (document → Markdown, OvisOCR2 :8017) |
MEETING_OCR_BASE_URL |
http://127.0.0.1:8017/v1 |
OCR backend endpoint |
MEETING_LLM_API_KEY |
(none) | LLM API key |
MEETING_LLM_MODEL |
(none — fail-closed) | LLM model name |
MEETING_VERIFY_BASE_URL / _API_KEY / _MODEL |
(none — uses main LLM) | Tiered second-pass verification endpoint |
MEETING_PROTOCOL_VERIFY |
true |
false disables second-pass protocol verification |
MEETING_PROTOCOL_CHUNK_OVERLAP_LINES |
overlap of chunking | Lines of overlap between protocol chunks |
MEETING_ROOT |
platform default (~/Meeting win / ~/meeting posix) |
Meetings storage root |
MEETING_TG_API_ID / MEETING_TG_API_HASH |
(none) | Telegram API credentials for voice/video ingest |
MEETING_TRANSCRIBE_MODEL |
auto (small / large-v3-turbo on CUDA) |
Default Whisper model |
MEETING_TRANSCRIBE_DEVICE |
auto | cpu or cuda |
MEETING_TRANSCRIBE_COMPUTE |
int8 / float16 (CUDA) |
Whisper compute type |
MEETING_TRANSCRIBE_LANG |
ru |
Default source language |
MEETING_TRANSLATE_BATCH_SIZE |
8 |
Lines per LLM translation batch |
MEETING_MAX_FILE_MB |
2048 |
Max input file size |
MEETING_MAX_DURATION_SEC |
7200 |
Max media duration |
MEETING_PROTOCOL_CHUNK_SIZE |
6000 |
Token threshold for protocol chunking |
MEETING_AGENT_MODE |
false |
If true, CLI prints agent-ready JSON |
MEETING_YT_PROXY |
socks5://127.0.0.1:12334 (YouTube only) |
Proxy for yt-dlp downloads; unset = direct |
Safety
- Cloud blocked by default. External endpoints are rejected unless
--allow-cloud(CLI) /allow_cloud=True(API) is supplied. There is no env override by design — cloud access is a per-run explicit decision. - Grounded protocols. Every decision and assignment requires a
source_quoteverified against the transcript. Unverifiable items are flagged with warnings; fabricated items cause validation failure. - No secret logging. API keys and credentials are never written to logs.
- Audit metadata. Each protocol includes
source_hash,stt_model,llm_model,created_at, andcloud_allowedfor traceability.
Development
pip install -e '.[dev]'
pytest -q # 202 tests (1 xfail host-specific)
The dev extra pulls everything the test suite imports (fastapi, mcp<2, openpyxl, telethon, PySocks, pytest). Verified on a clean Ubuntu 24.04 VM: without it the suite reports ~13 collection errors, with it — 201 passed / 1 skipped.
Architecture decisions: docs/adr/ · Design doc: docs/sdd-v0.8.0.md
License
MIT