CUSTOMER CASE STUDY
Powered byAmazon Web Services+Amazon BedrockAmazon Bedrock · Amazon EC2 · Amazon DocumentDB

Agentic Service Operations

Rebuilding after-sales operations around Amazon Bedrock agents — warnings before failures, dispatch before phone calls, parts before demand. Fleet unplanned downtime fell 41%.

Zhejiang EP Equipment Co., Ltd.

Client EP Equipment (EP · BigJoe)Coverage 100+ CountriesUse Case Generative AI · Agentic Operations
01
-41%
Unplanned Downtime
02
78%
First-Time Fix Rate
03
40s
Average Dispatch Time
04
47%
Preventive Work Orders
Business Challenges

In 2025 the company launched a Bedrock-based service-knowledge platform with AI-Driven. Answers got faster — but the operating model itself had not changed: break-fix service, manual dispatch, and judgment-based stocking.

01

Reactive Service, Idle Telemetry

Customers called only after breakdowns. Preventive work accounted for just 12% of maintenance orders, fleet unplanned downtime ran 14.6 hours per vehicle per half-year, and real-time telemetry from roughly 50,000 connected vehicles was only reviewed after the fact.

02

Manual Dispatch, Low First-Time Fix

Dispatchers hand-balanced skills, location, daily schedules, parts availability, and workload at about 25 minutes per order. Mis-assignment and missing parts kept the first-time fix rate at 61%, and repeat visits drove up both downtime and cost.

03

Coarse Global Parts Management

Eighteen overseas warehouses ran on three heterogeneous sources — a DAS system, an Australian WMS, and offline Excel. Stockouts hit 8.6%, triggering repeat visits and about $9,800 of emergency air freight per month, while multi-currency data and inconsistent part numbering blocked any global view.

04

Slow, Unauditable Settlement

Engineers spent about 20 minutes per order on paperwork with inconsistent fields. Reconciliation was manual, exceptions were caught by sampling, and nothing traced settlement back to the dispatch decision behind it.

Solution & Architecture

A two-layer "deterministic engine plus LLM agents" architecture with Amazon Bedrock at the core, where three business lines share one governed cloud foundation

Amazon Bedrock

The Amazon Bedrock Agent Families

Sonnet-tier Claude models drive the reasoning chains, a Haiku-tier model handles high-volume classification, routing, and pre-scoring, and Amazon Titan Text Embeddings V2 vectorizes knowledge and cases — all invoked privately through VPC endpoints, with enterprise data never used for model training

Dispatch Chain · 4 Stages

Fault understanding → repair planning → parts coverage → an LLM ranking agent that re-ranks the rule pipeline's top candidates and writes its rationale, fused 0.6/0.4 with deterministic scores

Proactive Service Agents

Daily autonomous fleet inspection and anomaly attribution, opening preventive work orders at least 48 hours before predicted failures

Parts & Strategy Agents

A replenishment agent reviews every MRP proposal and issues cited purchase proposals; daily agents tune dispatch parameters and a weekly agent evolves strategy from closed-loop data

L1
Amazon EC2 + Elastic Load Balancing

Proactive Service: Warnings Before Failures

TBox telemetry from roughly 50,000 vehicles enters through Elastic Load Balancing into self-managed stream processing and a complex-event engine on Amazon EC2 inside private VPC subnets, where trend analysis plus an expert-rule base raises risk warnings that Bedrock agents attribute and turn into work orders.

≤5s processing latency, ≥1,000 events per second

50+ scenario expert-rule base layered on trend analysis

Preventive orders opened 48+ hours ahead automatically

L2
Amazon Bedrock + Amazon DocumentDB

Agentic Dispatch & Settlement: Dispatch Before Calls

The Bedrock agent chain resolves fault features and parts coverage first; a deterministic dispatch pipeline on EC2 (parallel hydrators, hard filters, multi-dimension scorers, selector) produces candidates, and the LLM ranking agent issues the final order with its rationale. Settlement runs on AI pre-fill, scan-to-verify, and slider confirmation, reconciled by a dual rule-plus-LLM engine.

Technician profiles across 12+ dimensions, filtered and scored in parallel

Every decision fully traced in Amazon DocumentDB

AI pre-fill and scan-to-verify; 96% automated reconciliation

L3
Amazon RDS + AWS Service Catalog

Global Parts Intelligence: Parts Before Demand

Three data sources across 18 overseas warehouses — DAS, WMS, and offline spreadsheets — feed a daily ETL into Amazon RDS (PostgreSQL, Multi-AZ). A two-step MRP engine computes safety stock and replenishment proposals, and a Bedrock replenishment agent reviews each one against logistics lead times, promotions, and product transitions.

~54,000 inventory snapshot rows per day, unified on a USD base

Every proposal issued as a cited purchase recommendation

Service Catalog replicates environments per brand and region

Autonomy Rollout & Governance

Agents don't switch to autopilot overnight. A shadow → suggestion → autopilot rollout, backed by circuit breakers, full traceability, and hard safety limits, keeps every expansion of autonomy explainable, auditable, and reversible.

01

Shadow Mode

Agents decided alongside dispatchers without acting. Double-blind review of every disagreement surfaced undigitized tacit rules — contract-exclusive technicians, customer blacklists, restricted products — which were encoded into the pipeline's hard filters.

Disagreement fell from 34% to 9% before advancing
02

Suggestion Mode

Agents propose a ranked shortlist with natural-language rationale for dispatchers to confirm; every acceptance and override flows back as a training signal for strategy evolution.

84% suggestion-mode adoption
03

Autopilot Mode

Rollout is segmented across four dimensions — product SLA, customer tier, region, and order type — so management can reclaim autonomy along any one of them at any time.

Autopilot covers 62% of work orders
Guardrails
3-Second Circuit Breaker

Out-of-range parameters or malformed output degrade to suggestion mode with an alert — no silent failures

100% Decision Traces

Every candidate's elimination reason, dimension scores, and LLM rationale are retained in Amazon DocumentDB

Bedrock Guardrails

Live-electrical and lifting work is always routed to humans; customer-facing output is PII-masked

Private VPC Endpoints

Bedrock is reached from private subnets, and enterprise data is never used for model training

0.6 / 0.4 Score Fusion

LLM ranking is fused with deterministic scores rather than replacing them, preserving predictable fallback behaviour

Versioned, Revertible Rules

Weekly strategy evolution requires p<0.05 validation, backtesting, and a 10% canary before full rollout

Outcomes & Metrics

Measured January–June 2026 from the customer's work-order system, telemetry platform, dispatch decision logs, and global parts platform. All figures are measured, not projected.

Fleet Unplanned Downtime

Telemetry platform + work-order system
Baseline

14.6 hrs per vehicle per half-year

Result

8.6 hrs per vehicle per half-year

−41%

Preventive Work Orders

Work-order system (order-origin tagging)
Baseline

12%

Result

47%, triggered automatically by risk warnings

+35 pp

Failure Risk Warnings

Proactive service platform warning logs
Baseline

No system in place

Result

48+ hours ahead; 87% precision; 81% alert-handling rate

New capability

First-Time Fix Rate

Work-order system
Baseline

61%

Result

78%

+17 pp

Dispatch Handling Efficiency

Dispatch decision logs (100% traced)
Baseline

25 minutes per order, manual

Result

40 seconds per order in autopilot; P95 under 3 minutes

+210% productivity

Global Parts Stockout Rate

Daily snapshots across 18 warehouses + DAS/WMS
Baseline

8.6%

Result

5.2%, with inventory turns up 17%

−40%

Settlement & Reconciliation

Settlement system
Baseline

~20 minutes per order; manual reconciliation

Result

28 seconds per order (P90); 96% automated reconciliation

96% automated

Degree of Autonomy (Adoption)

Dispatch system operating statistics
Baseline

None

Result

84% suggestion-mode adoption; autopilot covers 62% of orders

48h+
Warning Lead Time

87% warning precision and an 81% alert-handling rate in canary

100%
Decisions Traced

Elimination reasons and LLM rationale remain fully auditable

18
Warehouses Unified

Three heterogeneous sources feed one governed platform daily

~50K
Connected Vehicles

TBox telemetry streaming continuously at ≤5s processing latency

TCO & Investment Analysis

Agent invocations grew about 2.8× as autopilot expanded, yet average model cost per work order fell 47%. As autonomy coverage widens and parts working capital declines, three-year total TCO is expected to run about 50% below a self-hosted-GPU-plus-headcount path (projected).

$4,800
Average Monthly Cloud Spend

Measured January–June 2026 via AWS Cost Explorer. LLM ranking runs only over the pipeline's top-5 candidates, Haiku-tier pre-scoring absorbs high-volume light tasks, and strategy agents batch off-peak — together holding unit costs down.

$208,800
Annualized Cost Savings

The cost pool of dispatch and settlement staffing plus emergency air freight fell from $40,700 to $23,300 per month including incremental AWS spend — $17,400 saved monthly, a ~43% reduction — as dispatch and reconciliation staff moved to exception handling and rule governance.

-89%
Inference Layer vs Self-Hosted

Bedrock spend averages about $2,500 per month against an estimated $22,000 per month for an equivalent self-managed GPU inference cluster plus 0.5 MLOps FTE.

Cost_Breakdown
Amazon Bedrock 52%
Amazon EC2 22%
Amazon RDS 10%
Amazon DocumentDB 9%
ELB / VPC & Others 7%
Next_Step

Let Agents Run Your Service Operations

See how AI-Driven builds explainable, auditable, and reversible agentic operations platforms on Amazon Bedrock