Kindo × Deloitte — High-Level Gantt with Dependencies

SOC Pillar (A.1–A.5) + New Agents (A.6, A.7, A.11) · June 9 – July 4, 2026 · 4 Weeks

Program Meeting — July 22, 2026

Biweekly program call · Platform improvements + agent capabilities
Attendees: Charlie Hulcher, Zun Huang, Nathan, Matthew Lew, Tony
Priority 1 Agent Observability & Failure Diagnostics
Production SOC agents run across two Deloitte clients. When an agent fails or stops mid-run, there is no reliable way to determine the root cause — logs roll over within ~20 minutes, and no structured failure data is captured. This blocks timely incident response and erodes confidence in agent-driven automation.
What Deloitte Needs
  • Durable logs that survive beyond 20-min rollover window
  • Structured failure reasons per agent run (tool error, timeout, context overflow)
  • Per-step and per-run token/cost tracking for usage attribution
  • Self-serve debug bundle export from failed runs (L1 triage without infra access)
  • Replicable monitoring dashboards for their own environment
What Kindo Has Today
  • Grafana Cloud — operational health (ingress, Hatchet task execution, run status)
  • Langfuse — LLM-level tracing (prompts, tool selection, model calls)
  • Audit logs — user-activity trail with Syslog/DB delivery + CSV export
  • Turnkey SMK auto-forwards to CloudWatch + X-Ray
  • ENG-10858 — Fix run status stuck "in_progress" on abrupt failure
Gaps & Open Items
  • Non-turnkey SMK (Deloitte's deployment) has no durable log backend — logs lost before investigation starts
  • No Grafana dashboard JSON or OTEL metric docs shared for self-hosted replication
  • No container-log forwarding documentation for non-turnkey deployments
  • No one-click debug bundle for failed agent runs
  • No per-step token breakdown or cost attribution
Desired Outcome from This Call

Align on which observability capabilities ship first. Determine whether Grafana dashboard export can be delivered as a quick win. Confirm timeline for debug bundle (ENG-10478) and log backend wiring (ENG-10512) — the two items marked high priority.

Priority 2 MCP Tool-Calling & Agent Step Latency
The SOC Level-1 triage agent takes 10–15 minutes per ticket on Kindo, compared to 3–8 minutes on the prior Swimlane-based approach. The Jun 29 review confirmed that prompt trimming and model selection do not materially reduce runtime — the primary driver is MCP tool-call retries where the agent selects an incorrect tool, fails, reasons about the failure, and retries with a different tool.
What Deloitte Needs
  • Triage agent completing in under 8 minutes per ticket
  • Reduced tool-call retry loops — agent picks the right tool on the first attempt
  • Visibility into which tool calls are failing and why (ties to observability)
  • Ability to constrain which tools are available per step
What Kindo Has Today
  • Multi-step agent builder with per-step model selection
  • MCP tool integration with LLM-driven tool selection
  • Conversation compaction for long-running sessions (recently shipped)
  • File attachments per step for deterministic context delivery
  • Run history showing step-level outputs and tool call attempts
Gaps & Open Items
  • No per-step tool scoping — agent sees all tools, increasing selection errors
  • No quantified retry-vs-productive-time breakdown per run
  • Inter-step invocation overhead not benchmarked (framework vs context passing)
  • Charlie's recommendations from Jun 29 review — status to confirm
Desired Outcome from This Call

Review Charlie's recommendations from the Jun 29 agent architecture session. Determine whether per-step tool scoping is feasible near-term. Set a realistic target latency and identify the split between platform improvements and prompt/architecture optimizations Deloitte can apply now.

Priority 2 Deterministic Flow Control & Automatic Retry
SOC agents require conditional routing (e.g., EPDR detection → EPDR skill file, XDR → XDR skill file) and automatic recovery from transient failures. Today, routing is handled implicitly through LLM reasoning rather than deterministic logic, and any step failure requires manual re-run. This limits agent reliability and increases operational overhead.
What Deloitte Needs
  • Conditional branching — "if detection type = EPDR, run these steps"
  • Loop / iteration — "for each alert in batch, run triage"
  • Configurable auto-retry per step on transient failure
  • Parallel step execution for independent operations
What Kindo Has Today
  • Linear multi-step agent execution with per-step configuration
  • Per-step model selection and file attachments
  • Re-run agent from a specific failed step (shipping next build)
  • Restart from previous run ID (shipping next build)
  • Agent import/export between installations (shipping next build)
Gaps & Open Items
  • No conditional branching between steps — classification → routing is LLM-implicit
  • No loop/iteration construct for batch processing
  • No configurable auto-retry with backoff per step
  • No parallel step execution
  • Product direction unclear: visual workflow builder vs enhanced step primitives
Desired Outcome from This Call

Understand Kindo's product direction for flow control. Determine whether auto-retry is a near-term deliverable. Identify interim architectural patterns Deloitte can use today (e.g., multi-agent chaining, classification-first routing) while platform capabilities mature.

Connecting Thread

These three topics are interconnected: MCP tool-call retries inflate latency (P2), but diagnosing which tools fail requires observability that doesn't exist yet (P1), and preventing retries entirely requires deterministic flow control that isn't available (P2). Solving any one accelerates the others. The program call should establish which capabilities are already in the Kindo pipeline, which require new product decisions, and which gaps can be mitigated with prompt and architecture changes available today.

TEK Team (connectors)
APO Team (API/MCP)
Kindo ENG (platform)
T&C (we own)
Gate / Milestone
▸ Pred = depends on
▸ Succ = feeds into
#
Task
Week 1
Jun 9–13
Week 2
Jun 16–20
Week 3
Jun 23–27
Week 4
Jun 30–Jul 4
TEK Integration Connectors — SOC Blockers
1
TEK-154 Swimlane 0-record fetch
Neha · URGENT
▸ Pred: none · ▸ Succ: 12, 15
✓ Done
·
2
TEK-95 VirusTotal enrichment
Harshal · URGENT · Since Apr 11
▸ Pred: none · ▸ Succ: 12, 15
✓ Done (next build)
·
3
TEK-106 ThreatConnect reliability
URGENT · Since Apr 23
▸ Pred: none · ▸ Succ: 12, 15
✓ Done (next build)
·
4
TEK-141 Swimlane record update
Harshal · TEST PROD
▸ Pred: none (ready) · ▸ Succ: 13
✓ Done (next build)
·
5
TEK-58 SailPoint ISC coverage
Deloitte · TEST PROD
▸ Pred: none (ready) · ▸ Succ: 13, 24
◆ Deloitte validation
·
6
TEK-139 SailPoint write ops
Sumanth · DONE
▸ Pred: Deloitte evidence · ▸ Succ: 12
✓ Done (next build)
·
7
TEK-155 SailPoint tool policies
Sumanth · DONE (next build)
▸ Pred: none · ▸ Succ: 12
✓ Done
·
8
TEK-128 Saviynt connector
Sumanth · DONE
▸ Pred: none · ▸ Succ: 12, 24
✓ Done (next build)
·
9
TEK-143 GitHub OAuth on SMK
Divya · BACKLOG
▸ Pred: none · ▸ Succ: 12 (A.4)
Backlog — deprioritized
APO API / MCP Connectors — SOC Blockers
10
APO-117 Jira upload attachments
Harshal · DONE (next build)
▸ Pred: none · ▸ Succ: 12, 19
✓ Done (next build)
·
11
APO-95 SQL Database MCP
Deloitte · IN REVIEW
▸ Pred: none · ▸ Succ: 12 (A.3, A.5)
In review
Target: ship
·
ENG Kindo Platform — Cross-Agent
12
Long-running reliability (85%→100%)
IN PROGRESS
▸ Pred: none · ▸ Succ: 14
Remaining 15%
Target: done
·
13
Multi-Agent Orchestration (61%)
IN PROGRESS
▸ Pred: none · ▸ Succ: 17 (A.6 chaining)
In progress
In progress
Target completion
·
◆ GATE A.1–A.5 Production Readiness
14
A.1–A.5 E2E Integration Testing
T&C coordinates
▸ Pred: 1,2,3,4,5,6,7,8,10,11,12 · ▸ Succ: 15
·
E2E test all 5 agents
·
15
◆ Production Handoff to Krishna
T&C + Krishna's team
▸ Pred: 14 · ▸ Succ: 17 (A.6 needs live data)
·
◆ Handoff + release notes
·
T&C New Agent Development — We Own This
16
A.6 Requirements (questionnaire + Krishna interview)
Warren + Tony · Option 3
▸ Pred: none · ▸ Succ: 17
D1 Quest → D2 Mtg
D2 Spec
·
17
A.6 Agent Build on Kindo
Warren
▸ Pred: 16, (15 for live data) · ▸ Succ: 18
·
D3-4 Build (NL design)
·
18
A.6 Test + Validate
Warren · Needs test data
▸ Pred: 17, test data from Krishna · ▸ Succ: 19
·
D5-6 Test vs MTTD/MTTR
·
19
◆ A.6 Review/QA with Krishna
Tony + Krishna
▸ Pred: 18 · ▸ Succ: production deploy
·
◆ D7 QA Review
·
20
A.7 Design Sprint with QA team
Tony/Joana · Option 1
▸ Pred: none · ▸ Succ: 21
Sched
D2-3 Design Sprint (90m)
·
21
A.7 Spec + Agent Build
Warren
▸ Pred: 20, (10 for Jira upload) · ▸ Succ: 22
·
D4 Spec → D5-6 Build
·
22
A.7 Test + Validate
Warren
▸ Pred: 21 · ▸ Succ: 23
·
D7-8 Test
·
23
◆ A.7 Review/QA with Krishna
Tony + Krishna
▸ Pred: 22 · ▸ Succ: production deploy
·
◆ D9 QA Review
·
24
A.11 SOPs + Doc Ingestion
Warren · Option 2
▸ Pred: none · ▸ Succ: 25
D1-2 Request SOPs
D3-5 Ingest + map
·
25
A.11 Ride-along + Reqs Matrix
Tony/Joana + Tim Corder team
▸ Pred: 24 · ▸ Succ: 26
·
D5 Ride-along
D6-7 Reqs matrix
·
26
A.11 Agent Build on Kindo
Warren
▸ Pred: 25, (5+8 for SailPoint/Saviynt) · ▸ Succ: 27
·
D8-10 Build (IAM workflows)
·
27
A.11 Test + Validate
Warren
▸ Pred: 26 · ▸ Succ: 28
·
D11-12 Test
·
28
◆ A.11 Review/QA with Tim + Krishna
Tony + Tim Corder + Krishna
▸ Pred: 27 · ▸ Succ: production deploy
·
◆ QA Review

Predecessor / Successor Dependency Matrix

#TaskTeamPredecessors (depends on)Successors (feeds into)
TEK Team — Integration Connectors
1TEK-154 Swimlane 0-record fetchTEKNone→ 14 (E2E Testing) · → 15 (Handoff)
2TEK-95 VirusTotal enrichmentTEKNone→ 14 (E2E Testing) · → 17 (A.6 build — enrichment data)
3TEK-106 ThreatConnect reliabilityTEKNone→ 14 (E2E Testing)
4TEK-141 Swimlane record updateTEKNone (ready)→ 13 (Deloitte validation) · → 14
5TEK-58 SailPoint ISC coverageTEKNone (ready)→ 14 · → 26 (A.11 build — IAM)
6TEK-139 SailPoint write ops ✓TEK← Deloitte evidence (resolved)→ 14 (A.1/A.4 remediation actions)
7TEK-155 SailPoint tool policiesTEKNone→ 14
8TEK-128 Saviynt connectorTEKNone→ 14 (A.5 CTEM) · → 26 (A.11 build — IAM)
9TEK-143 GitHub OAuth on SMKTEKNone→ 14 (A.4 detection rules)
APO Team — API / MCP
10APO-117 Jira upload attachmentsAPONone→ 14 · → 21 (A.7 audit evidence)
11APO-95 SQL Database MCPAPONone→ 14 (A.3 hunt, A.5 exposure)
Kindo ENG — Platform
12Long-running reliabilityENGNone→ 14 (agent stability)
13Multi-Agent OrchestrationENGNone→ 17 (A.6 agent chaining with A.1-A.5)
◆ Gates
14A.1–A.5 E2E Integration TestingT&C← 1,2,3,4,5,6,7,8,9,10,11,12→ 15
15◆ Production Handoff to KrishnaGATE← 14→ 17 (A.6 needs live agent data)
T&C — A.6 Vitals Dashboard (Option 3 · 7 days)
16A.6 Requirements (questionnaire + interview)T&CNone (start immediately)→ 17
17A.6 Agent BuildT&C← 16 · ← 15 (live data, soft dep)→ 18
18A.6 Test + ValidateT&C← 17 · ← test data from Krishna→ 19
19◆ A.6 QA Review with KrishnaGATE← 18→ Production deploy
T&C — A.7 Quality Audit (Option 1 · 9 days)
20A.7 Design Sprint with QA teamT&CNone (start immediately)→ 21
21A.7 Spec + BuildT&C← 20 · ← 10 (Jira upload, soft dep)→ 22
22A.7 Test + ValidateT&C← 21→ 23
23◆ A.7 QA Review with KrishnaGATE← 22→ Production deploy
T&C — A.11 Identity Agent (Option 2 · 12 days)
24A.11 SOPs + Doc IngestionT&CNone (start immediately)→ 25
25A.11 Ride-along + Reqs MatrixT&C← 24→ 26
26A.11 Agent BuildT&C← 25 · ← 5,8 (SailPoint/Saviynt, soft dep)→ 27
27A.11 Test + ValidateT&C← 26→ 28
28◆ A.11 QA Review with Tim + KrishnaGATE← 27→ Production deploy

Critical Paths

Longest dependency chains that determine delivery dates. Red = external team (not our control). Green = T&C (we control). Yellow = gate/milestone.

A.1–A.5 Production (longest path):
1 TEK-154 2 TEK-95 3 TEK-106 14 E2E Test 15 ◆ Handoff
TEK-95 and TEK-154 both done. Critical path for A.1–A.5 is now E2E testing (14) → handoff (15).
A.6 Vitals Dashboard:
16 Reqs 17 Build 18 Test 19 ◆ QA Review
Fully in our control. Soft dep on 15 (live data) — mitigated by building against available data first.
A.7 Quality Audit:
20 Design Sprint 21 Spec+Build 22 Test 23 ◆ QA Review
Fully in our control. Soft dep on 10 (Jira upload for evidence attachments).
A.11 Identity Agent (longest T&C path):
24 SOPs 25 Ride-along 26 Build 27 Test 28 ◆ QA Review
Mostly in our control. Hard dep on 5 (SailPoint ISC) + 8 (Saviynt) for build step — both TEK team, both done.