Steel price report PDFs in; charts, spreads and an AI market read out, where every number is traced to a deterministic parser.
Turns weekly commodity price report PDFs into charts, spreads and a grounded AI market read. Numbers come from a deterministic parser; the LLM only explains them.
Highlights
- Grounded analyzer: verdict cards per market, top movers, spreads, risks and data gaps, with a self-check strip for numbers traced, units bound, NA handled as gaps.
- Hallucination register: 13 named failure modes (
R1-R14, noR12) written into the prompt pack. - Multi-agent pipeline: section map and price pack run in parallel, then the analyzer; results cached by content hash and prompt version.
- Degrades cleanly: no API key means a valid prices-only analysis, not an error.
- Bilingual chat: asks and answers in English or Persian (RTL) over the same price context.
Screenshots
All screens use the synthetic 2030 demo prices from backend/scripts/seed_demo.py. The analyzer and chat output shown is mocked (demo data), since those need an LLM key.
Overview: core watchlist rebased to 100 and the day’s movers. | Charts: any series in index, % or absolute mode, CSV export. |
Spreads: rebar–scrap margin and grade premiums, each on its own axis. | Analyzer: overall read and per-market verdict cards with source symbols (demo data). |
Analyzer details: movers with cadence and “so what”, computed spreads with formulas (demo data). | Chat: English and Persian questions over the price pack (demo data). |
Mobile: the analysis view at 390 px wide (demo data). |
Why it’s interesting
- Numbers and narrative are walled off. A deterministic price pack (SQLite) is the only source of numbers; the LLM may only add drivers and tone from the report text. Anything it cannot trace becomes a listed data gap.
- Hallucination-register prompt pack. 13 failure modes (
R1-R14) (cross-unit assignment, NA turned into 0, stale-date “daily” moves, range endpoints, double-counted converted rows) are written into the analyzer prompt and a self-check the model must fill in. - Strict, chart-ready output contract. The analyzer emits one JSON object (verdict cards, movers, spreads, margins, curve signals, risks, data gaps,
self_check) followed by a markdown brief, and the UI renders it literally. - Works without an LLM. With no API key, the pipeline still produces a valid analysis object from the price pack alone (
_deterministic_analysis). - Domain units handled explicitly.
$/dmt,$/dmtu,$/st,$/mt,Eur/mtand differing cadences (DoD, WoW, contract) are part of the prompt and the price pack.
All sample values in the prompt pack and tests are synthetic.
Architecture
flowchart LR
PDF[Price report PDFs] --> P[parse_pdf.py + catalog.yaml]
P --> DB[(SQLite)]
DB --> B[Agent B: deterministic price pack]
PDF --> S[sections.py: narrative sections]
S --> A[Agent A: section map]
B --> E[Agent E: grounded analyzer]
A --> E
S --> E
E --> UI[React UI: verdict cards, movers, spreads]
DB --> C[Chat orchestrator]
E --> C
C --> UI
Agents A and B run in parallel, then the analyzer runs last. The chat orchestrator answers from the same price context plus stored analyses.
Tech stack
FastAPI, SQLite, PyMuPDF, pandas, PyYAML, OpenRouter (default model google/gemini-3.5-flash-lite), React + Vite + TypeScript.
Key techniques
- Deterministic PDF-to-price extraction with a symbol catalog:
backend/app/services/parse_pdf.py,backend/app/services/catalog.yaml - Price pack (movers, spreads, gaps):
backend/app/agents/prices_agent.py - Grounded analyzer and no-LLM fallback:
backend/app/agents/analyzer_agent.py - Chat orchestrator:
backend/app/agents/orchestrator.py - Prompt pack (system prompts, few-shots, domain reference):
backend/app/prompts/ - Content-hash cache of AI results keyed by prompt version:
backend/app/services/store.py
The parser and catalog contain a few literal table-title anchors from the source report layout; adapt them to your own report format.
Getting started
python -m pip install -r backend/requirements.txt
cp .env.example .env # set OPENROUTER_API_KEY (optional; no key = no-LLM mode)
python -m uvicorn backend.app.main:app --reload --port 8000
cd frontend && npm install && npm run dev
UI: http://localhost:5173, API docs: http://127.0.0.1:8000/docs.
No PDFs? Seed the database with synthetic, fictional prices (18 issues dated 2030) to try the UI:
python backend/scripts/seed_demo.py
Regenerate the README screenshots with node scripts/capture-screenshots.cjs (needs Playwright; see the header comment for the ports it expects).
First use: Upload PDF page (parse PDFs from METAP_PDF_DIR, files named SD_YYYYMMDD.pdf), then Charts / Spreads, then AI Analysis, then Chat. You supply your own licensed report PDFs; none are included.
Tests
cd backend && python -m pytest tests -q
The test builds a temp SQLite DB with fictional prices and checks the deterministic price pack (movers, spreads, gaps, NA handling).
License
MIT, see LICENSE. Source reports are third-party copyrighted material and are not included; use only licensed copies.
Built by Sepehr Radmard · LinkedIn · GitHub · more projects on my profile
Overview: core watchlist rebased to 100 and the day’s movers.
Charts: any series in index, % or absolute mode, CSV export.
Spreads: rebar–scrap margin and grade premiums, each on its own axis.
Analyzer: overall read and per-market verdict cards with source symbols (demo data).
Analyzer details: movers with cadence and “so what”, computed spreads with formulas (demo data).
Chat: English and Persian questions over the price pack (demo data).
Mobile: the analysis view at 390 px wide (demo data).