Skip to content
All projects

03Multi-Agent System

Construction Plan AI Generation & Review System

建筑施工方案 AI 生成与审核系统

Role
Agent workflow design and implementation, retrieval pipeline and evaluation, plus deterministic tool-layer and execution-boundary design.
Stack
Multi-Agent · LangGraph · RAG · Tool Calling
Proof
0.60 Hit@1 · 1.00 Hit@3 · 0.77 MRR · 7 tests passed

01

Problem

01.1Overview

A multi-stage AI system for construction-plan generation and review, combining retrieval, deterministic tools and agent workflow orchestration.

The system assists with drafting and reviewing construction plans. It does not replace an engineer's professional judgment.

01.2Context

Writing a construction plan is spread across separate steps: looking up specifications, drafting the plan, estimating cost and duration, and reviewing the result. This system brings them together, combining RAG, tool calling and a multi-stage agent workflow.

01.3Problem

Specification lookup, plan drafting, cost / duration estimation and review happen in separate steps with little connection between them.

Handing the whole task to a single LLM call makes the grounding hard to trace — and calculations such as cost and duration need stable results that the model should not generate directly.

01.4My Role

Agent workflow design and implementation, retrieval pipeline and evaluation, plus deterministic tool-layer and execution-boundary design.

02

Build

02.1Solution

LangGraph StateGraph coordinates intent recognition, retrieval, generation, tool execution and review as separate stages instead of relying on a single LLM call.

Specification content is retrieved from a knowledge base first: PDF / DOCX documents are parsed and chunked, embedded with BGE-M3 (1024 dimensions) and searched with FAISS to give generation its context.

For tasks that require deterministic outputs, such as cost or duration estimation, the agent identifies intent and prepares parameters while registered tools perform the actual calculation.

Tool execution is constrained by permission checks, structured outputs and fail-closed handling so model decisions do not directly bypass the execution boundary.

02.2Architecture

  1. User Request
  2. Intent Recognition
  3. LangGraph StateGraph
  4. Knowledge Retrieval (RAG)
  5. Plan Generation
  6. Tool Decision
  7. Registry-based Tool Layer
  8. Deterministic Calculation
  9. Result Review
  10. Structured Output
fig. 03a — multi-stage agent workflow

The LLM does not do everything: retrieval supplies the grounding, the model interprets and drafts, deterministic calculation goes to tools, and the result is reviewed before output.

Retrieval pipeline: PDF / DOCX parsing → chunking → BGE-M3 embedding (1024-dimensional) → FAISS retrieval, with retrieval quality measured on offline samples.

03Drafting a construction plan
  1. Intent

    Intent recognised: plan generation

  2. StateGraph

    StateGraph routes to the matching stages

  3. Retrieval

    BGE-M3 + FAISS retrieve relevant specifications

  4. Generation

    Draft the plan from the retrieved context

  5. Review

    Review the result before output

  6. Output

    Structured output

One run of the pipeline, step by step.

02.3Key Decisions

  1. 01

    Multi-stage Workflow

    Retrieval, generation, tool execution and review run as separate stages instead of one LLM call doing everything.

  2. 02

    Retrieval Before Generation

    Answers that depend on construction specifications retrieve relevant content from the knowledge base first, so generation has context.

  3. 03

    Tools for Deterministic Tasks

    Cost, duration and other tasks that need stable results are not generated by the model; they run through the tool layer.

  4. 04

    Controlled Execution Boundary

    Permission validation, structured output and fail-closed handling limit what the agent can execute.

03

Engineering

03.1Engineering

  • LangGraph StateGraph
  • Multi-stage agent workflow
  • RAG
  • PDF / DOCX parsing and chunking
  • BGE-M3 embedding (1024-dimensional)
  • FAISS retrieval
  • Retrieval evaluation: Hit@K / MRR
  • Registry-based tool layer
  • Tool calling
  • Permission validation
  • Structured output
  • fail-closed handling
  • Gradio interface

03.2Results

Hit@1 · 10 offline samples
0.60
Hit@3 · 10 offline samples
1.00
MRR · 10 offline samples
0.77
tests passed · pytest
7

Retrieval was evaluated on 10 offline evaluation samples: Hit@1 = 0.60, Hit@3 = 1.00, MRR = 0.77.

pytest: 7 passed. End-to-end validated through the Gradio interface.

04

Notes

04.1What I Learned

Once the task is split into stages, each step can be checked and tested on its own, and a failure is easier to locate.

Deciding what the model should not do matters as much as the prompt.

04.2Tech Stack

  • Python
  • LangGraph
  • LangChain
  • DeepSeek-V3
  • RAG
  • BGE-M3
  • FAISS
  • Tool Calling
  • Gradio