- Context windows (contexts/*.yaml) scope answers to a defined domain - Optional DuckDuckGo web search behind ai.web_search.enabled (default off) - Stream partial answers into the overlay at first-token time - Default Whisper to local base.en (~9x faster); offline model loading - Priority-ordered loopback detection (BlackHole > Teams device) - Overlay: drag interior to move, edges to resize - Stop tracking model binaries (models/ is gitignored) - README/CLAUDE.md overhaul + tracked config.example.yaml Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
168 lines
9.2 KiB
Markdown
168 lines
9.2 KiB
Markdown
# Meeting Assistant
|
||
|
||
A real-time, on-device AI copilot for meetings. You **hold a key to talk** (or
|
||
toggle hands-free auto-listen), it detects questions (even ones phrased as plain
|
||
statements), and **streams answers live** into an overlay that is hidden from
|
||
screen-capture. It can also **read a question off your screen** when you drag a
|
||
box around it, and scope every answer to a **context window** you define (e.g.
|
||
"Python/Flask senior backend interview").
|
||
|
||
Everything runs locally: **faster-whisper** (`base.en` by default, fully offline)
|
||
for speech-to-text and a quantized **Qwen2.5-VL-7B** vision model (via
|
||
`llama.cpp`) that answers both spoken and on-screen questions. No audio or text
|
||
leaves your machine — unless you explicitly enable the optional web search.
|
||
|
||
## What it does
|
||
|
||
- **Push-to-talk or auto-listen.** Hold **Right Option (⌥)** to capture; release
|
||
to answer. Or press **Ctrl+Shift+M** to toggle hands-free auto-listen
|
||
(continuous VAD). Captures your microphone *and* the other participants' audio
|
||
(system output via a loopback device), tagged by speaker (`You` / `Them`).
|
||
- **Streams answers as they generate.** First words appear in the overlay at
|
||
first-token time (~1–3 s) with a live `▌` cursor, instead of waiting for the
|
||
full answer. The console prints latency metrics (`⏱ whisper`, `⏱ gen`,
|
||
`⚡ first words`) so you can see where time goes.
|
||
- **Context windows keep answers on-topic.** Define a scope in
|
||
`contexts/*.yaml` (e.g. a Flask interview). Every question is interpreted
|
||
inside that scope first — ask "what is STOMP" and you get the Flask-relevant
|
||
meaning first, then other variations. See [contexts/README.md](contexts/README.md).
|
||
- **Reads on-screen questions.** Press **Ctrl+Shift+Space**, drag a box over a
|
||
question (e.g. on a shared slide), and the vision model reads and answers it —
|
||
including multiple-choice, code, and math.
|
||
- **Detects real questions** — interrogatives, yes/no questions, requests
|
||
("explain the rollout plan"), embedded/plain-sentence questions
|
||
("I was wondering about the budget"), tag questions ("…, right?"), math, and
|
||
fact-shaped fragments ("difference between TCP and UDP").
|
||
- **Auto-fills the obvious.** Arithmetic is solved instantly (no LLM); short
|
||
factual lookups get a direct answer. Open-ended questions get a concise
|
||
**suggested reply** you can read out, using the live meeting transcript as context.
|
||
- **Movable, resizable overlay.** Drag the middle of the overlay to move it;
|
||
drag any edge or corner to resize. Press **Ctrl+Shift+H** to show/hide, and
|
||
**Ctrl+Shift+S** for a meeting summary so far.
|
||
- **Optional web search (off by default).** Ground factual answers in current
|
||
information via DuckDuckGo — see Configuration. Enabling it sends the question
|
||
text off-device.
|
||
|
||
## Hearing other participants (one-time macOS setup)
|
||
|
||
Your microphone only captures *you*. To also capture what the other participants
|
||
say, the assistant reads your system audio output through a virtual loopback
|
||
device. Devices are auto-detected in priority order
|
||
(`audio.loopback_keywords`): **BlackHole / Soundflower / VB-Cable / Loopback**,
|
||
then Microsoft Teams' own virtual device as a last resort.
|
||
|
||
- **BlackHole (recommended — works for Google Meet, Zoom, Teams, anything):**
|
||
it mirrors *all* system audio. Install with `brew install blackhole-2ch`
|
||
(needs your admin password; reboot or restart `coreaudiod` afterwards).
|
||
- **"Microsoft Teams Audio" (zero-install experiment, Teams-only):** Teams
|
||
installs this device itself and the app will use it if no true loopback
|
||
exists — but it only ever carries Teams audio, and may be silent outside
|
||
screen-share. It will never help for Google Meet or Zoom.
|
||
|
||
The catch with BlackHole: if you send audio *only* to BlackHole, you won't hear
|
||
it yourself. So create a **Multi-Output Device** that plays to both your
|
||
speakers/headphones and BlackHole at once:
|
||
|
||
1. Open **Audio MIDI Setup** (Applications → Utilities).
|
||
2. Click **+** (bottom-left) → **Create Multi-Output Device**.
|
||
3. Check both **BlackHole 2ch** and your normal output (e.g. *MacBook Pro Speakers*).
|
||
4. In **System Settings → Sound → Output**, select that Multi-Output Device.
|
||
5. In your meeting app (Zoom/Teams/Meet), make sure the speaker/output is the
|
||
Multi-Output Device (or the system default).
|
||
|
||
Now meeting audio reaches both your ears and the assistant. Your microphone stays
|
||
selected as the meeting's *input*.
|
||
|
||
## Configuration
|
||
|
||
`config.yaml` is gitignored (it may hold machine-specific/private settings) —
|
||
copy `config.example.yaml` to `config.yaml` and edit. Key options:
|
||
|
||
```yaml
|
||
audio:
|
||
source: "both" # "microphone", "system", or "both"
|
||
capture_mode: "push_to_talk" # push_to_talk (hold key) or continuous (auto-listen)
|
||
ptt_key: "alt_r" # push-to-talk key (Right Option). e.g. cmd_r, f8, ctrl_r
|
||
# base.en is ~9x faster than medium.en on CPU (~0.3s vs ~2.9s per clip) and is
|
||
# the main latency fix. Use models/whisper-medium.en for noisy audio/accents.
|
||
whisper_model: models/whisper-base.en
|
||
mic_device: null # null = auto. Or a device index / name substring.
|
||
system_device: null # null = auto-detect loopback. Or index / name.
|
||
loopback_keywords: ["blackhole", "soundflower", "vb-cable", "loopback", "teams audio"]
|
||
answer_sources: # which speakers trigger an answer
|
||
- "system" # other participants
|
||
- "microphone" # your own voice (handy for testing)
|
||
# Incremental transcription while the key is held. Leave off with fast models:
|
||
# faster-whisper pads every call to a fixed 30s window, so it only pays off
|
||
# with large/slow models — and can drop words at chunk seams.
|
||
streaming_transcription: false
|
||
|
||
screen:
|
||
capture_key: "ctrl+shift+space" # press, then drag a box over a question
|
||
|
||
ai:
|
||
model: "Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf"
|
||
mmproj: "mmproj-Qwen2.5-VL-7B-Instruct-f16.gguf" # required for screen reading
|
||
answer_mode: "auto_obvious" # auto_obvious | auto_all | suggest_only
|
||
user_name: "you"
|
||
# active_context: python_flask_backend # pin a context file by name (optional)
|
||
# PRIVACY: web search sends the question text to DuckDuckGo. Off by default.
|
||
web_search:
|
||
enabled: false
|
||
max_results: 3
|
||
timeout: 6.0
|
||
```
|
||
|
||
- `auto_obvious` — answer obvious questions (math/factual) directly; show a
|
||
*suggested reply* for open-ended ones.
|
||
- `auto_all` — generate a full answer for every detected question.
|
||
- `suggest_only` — never commit; always show a draft.
|
||
|
||
### Context windows (`contexts/`)
|
||
|
||
Each `.yaml` file defines a named scope the AI reads before answering, so it
|
||
stays on-topic without you restating the context in every question. Mark one
|
||
`active: true` (or set `ai.active_context`). Files hot-reload — edit mid-meeting
|
||
and it takes effect on the next question. `strict: true` disables the
|
||
"…then note other variations" behavior. A complete example ships in
|
||
[`contexts/python_flask_backend.yaml`](contexts/python_flask_backend.yaml).
|
||
|
||
## Run
|
||
|
||
```bash
|
||
./run.sh # macOS / Linux (handles venv, CA bundle, PortAudio)
|
||
# or
|
||
python main.py
|
||
```
|
||
|
||
First run downloads the models (run `python setup.py` once, ~6 GB for the vision
|
||
model). The Whisper models load from local `models/whisper-*` directories, fully
|
||
offline — no Hugging Face access needed at runtime (works behind SSL-inspecting
|
||
corporate proxies).
|
||
|
||
On first launch grant three macOS permissions (System Settings → Privacy &
|
||
Security) **to the app you launch from** (Terminal, iTerm, or PyCharm):
|
||
- **Microphone** — to capture audio.
|
||
- **Screen & System Audio Recording** — without it, screen grabs silently
|
||
capture only your wallpaper (macOS returns a windowless desktop).
|
||
- **Accessibility** — without it, the global hotkeys (push-to-talk, screen grab)
|
||
never fire; the log shows `This process is not trusted!`.
|
||
|
||
Quit and relaunch after granting — macOS applies these on restart.
|
||
|
||
You can also type a question in the terminal + Enter to test the AI directly.
|
||
|
||
## How it works
|
||
|
||
| File | Role |
|
||
|------|------|
|
||
| `src/audio_listener.py` | Captures each source. Push-to-talk buffers audio only while armed; transcribes with one shared Whisper model (serialized — not thread-safe), tags by source. Priority-ordered loopback detection. Optional incremental (streaming) transcription for slow models. |
|
||
| `src/hotkeys.py` | Global hotkeys (pynput): hold-to-talk, screen-grab chord, overlay toggle, summary, auto-listen toggle. |
|
||
| `main.py` | Question detection + orchestration; wires streamed partial answers into the overlay with latency metrics. |
|
||
| `src/ai_engine.py` | Math fast-path, classification, context-window injection, optional web search, and Qwen2.5-VL answering (streamed) for spoken and on-screen (`answer_from_image`) questions. Interrupts only in-flight generations so rapid follow-ups are never dropped. |
|
||
| `src/context_library.py` | Loads `contexts/*.yaml`, picks the active context, hot-reloads on file change. |
|
||
| `src/web_search.py` | Optional dependency-free DuckDuckGo search (Instant Answer + HTML results), proxy-tolerant. |
|
||
| `src/region_capture.py` | Draw-a-box fullscreen selector + Retina-correct region screenshot (freeze-frame at hotkey time). |
|
||
| `src/context_manager.py` | Rolling speaker-labeled meeting transcript. |
|
||
| `src/overlay.py` | Always-on-top overlay, hidden from screen capture. Drag interior to move, edges/corners to resize. |
|