- Agent workflow design and implementation, retrieval pipeline and evaluation, plus deterministic tool-layer and execution-boundary design.
- Multi-Agent · LangGraph · RAG · Tool Calling
- 0.60 Hit@1 · 1.00 Hit@3 · 0.77 MRR · 7 tests passed
Problem
A multi-stage AI system for construction-plan generation and review, combining retrieval, deterministic tools and agent workflow orchestration.
The system assists with drafting and reviewing construction plans. It does not replace an engineer's professional judgment.
Writing a construction plan is spread across separate steps: looking up specifications, drafting the plan, estimating cost and duration, and reviewing the result. This system brings them together, combining RAG, tool calling and a multi-stage agent workflow.
Specification lookup, plan drafting, cost / duration estimation and review happen in separate steps with little connection between them.
Handing the whole task to a single LLM call makes the grounding hard to trace — and calculations such as cost and duration need stable results that the model should not generate directly.
Agent workflow design and implementation, retrieval pipeline and evaluation, plus deterministic tool-layer and execution-boundary design.
Build
LangGraph StateGraph coordinates intent recognition, retrieval, generation, tool execution and review as separate stages instead of relying on a single LLM call.
Specification content is retrieved from a knowledge base first: PDF / DOCX documents are parsed and chunked, embedded with BGE-M3 (1024 dimensions) and searched with FAISS to give generation its context.
For tasks that require deterministic outputs, such as cost or duration estimation, the agent identifies intent and prepares parameters while registered tools perform the actual calculation.
Tool execution is constrained by permission checks, structured outputs and fail-closed handling so model decisions do not directly bypass the execution boundary.
- User Request
- Intent Recognition
- LangGraph StateGraph
- Knowledge Retrieval (RAG)
- Plan Generation
- Tool Decision
- Registry-based Tool Layer
- Deterministic Calculation
- Result Review
- Structured Output
The LLM does not do everything: retrieval supplies the grounding, the model interprets and drafts, deterministic calculation goes to tools, and the result is reviewed before output.
Retrieval pipeline: PDF / DOCX parsing → chunking → BGE-M3 embedding (1024-dimensional) → FAISS retrieval, with retrieval quality measured on offline samples.
- Intent
- StateGraph
- Retrieval
- Generation
- Review
- Output
- 03› BGE-M3 + FAISS retrieve relevant specifications
- 04› Draft the plan from the retrieved context
- 05› Review the result before output
- 06› Structured output
Intent
Intent recognised: plan generation
StateGraph
StateGraph routes to the matching stages
Retrieval
BGE-M3 + FAISS retrieve relevant specifications
Generation
Draft the plan from the retrieved context
Review
Review the result before output
Output
Structured output
- 01
Multi-stage Workflow
Retrieval, generation, tool execution and review run as separate stages instead of one LLM call doing everything.
- 02
Retrieval Before Generation
Answers that depend on construction specifications retrieve relevant content from the knowledge base first, so generation has context.
- 03
Tools for Deterministic Tasks
Cost, duration and other tasks that need stable results are not generated by the model; they run through the tool layer.
- 04
Controlled Execution Boundary
Permission validation, structured output and fail-closed handling limit what the agent can execute.
Engineering
- LangGraph StateGraph
- Multi-stage agent workflow
- RAG
- PDF / DOCX parsing and chunking
- BGE-M3 embedding (1024-dimensional)
- FAISS retrieval
- Retrieval evaluation: Hit@K / MRR
- Registry-based tool layer
- Tool calling
- Permission validation
- Structured output
- fail-closed handling
- Gradio interface
- Hit@1 · 10 offline samples
- 0.60
- Hit@3 · 10 offline samples
- 1.00
- MRR · 10 offline samples
- 0.77
- tests passed · pytest
- 7
Retrieval was evaluated on 10 offline evaluation samples: Hit@1 = 0.60, Hit@3 = 1.00, MRR = 0.77.
pytest: 7 passed. End-to-end validated through the Gradio interface.
Notes
Once the task is split into stages, each step can be checked and tested on its own, and a failure is easier to locate.
Deciding what the model should not do matters as much as the prompt.
- Python
- LangGraph
- LangChain
- DeepSeek-V3
- RAG
- BGE-M3
- FAISS
- Tool Calling
- Gradio