AI Governance track · Platform engineering

AI Platform Engineering for Investment Management

Build the platform. Control the agents. Defend the architecture.

A production engineering programme for lead engineers, architects and senior data and ML engineers building enterprise AI platforms inside hedge funds, asset managers and investment banks. The centre of gravity is distributed systems, reusable platform capabilities, agent control, evaluation, security, observability and model risk, not prompt engineering.

Advanced practitioner to technical lead · Python first · concept brief, architecture clinic, implementation lab, failure exercise, design defence

Enrol or enquire See the 120 chapters

120 chapters · 12 modules · 24 engineering labs · 12 design reviews · one cumulative capstone

The cumulative case study

Aegis Investment Intelligence Fabric

The cohort builds one platform throughout the programme, serving portfolio managers, quantitative researchers, fundamental analysts, risk, compliance and platform engineers. Every module extends the same system rather than starting a new exercise.

What it ends up doing

Research ingestion across filings, transcripts, internal notes and structured market data; point-in-time, permission-aware retrieval with citations; issuer and sector research agents; portfolio exposure and scenario tools; investment-memo assembly with contradiction and evidence checks; research-code generation in isolated execution environments.

And how it is run

Multi-agent orchestration with budgets, approvals and durable state; offline and online evaluation, tracing, cost and latency controls; Kubernetes deployment, GitOps promotion and production incident response.

The programme does not claim that an autonomous agent should make unsupervised investment decisions. Portfolio decisions, trade instructions and external communications remain subject to explicit human authority, policy and auditable controls. Aegis is a fictional firm created for teaching.

Target architecture

Nine layers, each with an owner and a contract

The programme treats the platform as a set of reusable services rather than one application, so each layer can be evolved, tested and governed independently.

LayerResponsibilities Candidate technologies
ExperienceAnalyst workspace, APIs, streaming responses, approvalsFastAPI, Streamlit or React
Agent runtimeState machines, routing, tool execution, durable workflowsLangGraph/LangChain, LlamaIndex, custom Python
Model gatewayProvider abstraction, routing, quotas, caching, policyLiteLLM-compatible gateway or custom service
KnowledgeParsing, indexing, hybrid retrieval, reranking, provenancePostgreSQL/pgvector or Qdrant, object storage
DataMarket, fundamental, portfolio and research pipelinesSQL, Polars, Pandas, dbt, Airflow or Dagster
EvaluationDatasets, judges, deterministic checks, experiment lineageMLflow plus a custom evaluation harness
PlatformContainers, orchestration, deployment and secretsDocker, Kubernetes, Helm, Flux
OperationsLogs, metrics, traces, alerts, SLOs and cost attributionOpenTelemetry, Prometheus, Grafana, Loki/Tempo
GovernanceEntitlements, audit, data lineage, model risk, change controlOPA-style policy, IAM, evidence registers
Learning outcomes

What a graduate can do

  1. Translate an investment workflow into a bounded AI product with measurable outcomes
  2. Choose models, retrieval strategies, tools and orchestration patterns using explicit trade-offs
  3. Design multi-tenant, permission-aware AI platform services
  4. Build advanced retrieval across narrative, tabular, time-series and point-in-time data
  5. Design stateful, recoverable agents and multi-agent workflows
  6. Evaluate task quality, groundedness, safety, latency, cost and business impact
  7. Deploy and operate AI services on Kubernetes using GitOps
  8. Implement model, data and agent governance suitable for financial institutions
Curriculum

12 modules · 120 chapters

Each module carries its engineering labs and a design review, because the test of platform work is whether the decisions survive challenge.

M1 Investment AI Systems and Product Framing Ch 1-10
  1. Why investment AI is a systems problem
  2. Investment lifecycle: idea, research, sizing, execution and monitoring
  3. Users: PMs, analysts, quants, risk, operations and compliance
  4. Decision support versus automated decision authority
  5. Selecting high-value workflows
  6. Commercial-impact trees and baseline measurement
  7. Functional and non-functional requirements
  8. Build, buy and platform-boundary decisions
  9. Reference architecture for Aegis
  10. Architecture decision records and technical leadership

Labs Map an issuer-research workflow; write the platform charter and initial ADR set.

Design review Defend scope, control boundaries, SLOs and expected value.

M2 LLM Engineering Foundations at Production Depth Ch 11-20
  1. Transformer behaviour engineers must understand
  2. Tokenisation, context limits and information loss
  3. Model families, deployment modes and licensing
  4. Prompt contracts, structured output and schema validation
  5. Tool calling and constrained generation
  6. Model routing by risk, quality, latency and cost
  7. Context caching, batching and response streaming
  8. Retry, timeout, fallback and circuit-breaker design
  9. Model gateways, quotas and tenant isolation
  10. Reproducibility under model and prompt change

Labs Build a typed model gateway; benchmark three routing policies under failure and load.

Design review Approve the model abstraction and failure semantics.

M3 Investment Data Foundation Ch 21-30
  1. Data classes: market, fundamental, alternative, portfolio and research
  2. Batch, streaming and event-driven ingestion
  3. Point-in-time correctness and look-ahead bias
  4. Corporate actions, identifiers and entity resolution
  5. Document parsing and layout preservation
  6. Tabular extraction and validation
  7. Pandas versus Polars execution trade-offs
  8. SQL models, tests and contracts with dbt
  9. Airflow versus Dagster orchestration choices
  10. Lineage, provenance, licensing and retention

Labs Create a point-in-time issuer data mart; build a tested document-to-object-store pipeline.

Design review Demonstrate that the dataset cannot leak future information.

M4 Retrieval Beyond Basic RAG Ch 31-40
  1. Retrieval failure taxonomy
  2. Chunking by document semantics and layout
  3. Embeddings and domain-specific retrieval behaviour
  4. Vector index design and filtering
  5. Lexical, semantic and hybrid retrieval
  6. Query decomposition, expansion and rewriting
  7. Cross-encoder and LLM reranking
  8. Parent-child, graph and hierarchical retrieval
  9. Permission-aware and point-in-time retrieval
  10. Citation integrity, freshness and retrieval evaluation

Labs Implement hybrid financial-document retrieval; build adversarial retrieval tests.

Design review Compare retrieval designs using recall, precision, attribution and latency.

M5 Tools, Code Execution and Agent Runtime Ch 41-50
  1. Agent loop anatomy and stopping conditions
  2. Tool interface design and semantic contracts
  3. Read tools versus side-effecting tools
  4. Idempotency, retries and transaction boundaries
  5. Sandboxed Python and research-code execution
  6. SQL generation with semantic and policy guards
  7. Memory: working, episodic and durable state
  8. Planning patterns and when planning degrades performance
  9. Human approval and interruptible execution
  10. Durable checkpoints, replay and recovery

Labs Build an issuer-analysis agent with six typed tools; safely execute generated analytics code.

Design review Threat-model every tool and approve authority boundaries.

M6 Multi-Agent Architecture and Orchestration Ch 51-60
  1. When multiple agents are justified
  2. Supervisor, router, graph and blackboard patterns
  3. Role and capability decomposition
  4. Shared state versus isolated context
  5. Delegation protocols and task contracts
  6. Parallelism, joins and partial failure
  7. Consensus, debate and evidence aggregation
  8. Preventing loops, deadlocks and agent collusion
  9. Token, time and tool budgets
  10. Deterministic orchestration around probabilistic components

Labs Build research, risk and critic agents; implement bounded parallel orchestration with recovery.

Design review Prove that multi-agent complexity beats a simpler baseline.

M7 Investment Workflows and Quantitative Integration Ch 61-70
  1. Earnings and filing intelligence
  2. Transcript change and management-tone analysis
  3. News-event triage and entity linking
  4. Comparable-company and sector analysis
  5. Portfolio exposure and concentration tools
  6. Scenario construction and factor shocks
  7. Natural-language-to-research-code workflows
  8. Investment-memo generation with claim verification
  9. Pre-trade decision support and restricted actions
  10. Post-decision monitoring and thesis drift

Labs Implement an evidence-linked investment memo; build a portfolio scenario agent with deterministic calculations.

Design review PM and risk personas challenge the workflow’s usefulness and controls.

M8 Evaluation and Experimental Discipline Ch 71-80
  1. Evaluation architecture and quality taxonomy
  2. Golden datasets and expert annotation
  3. Deterministic, statistical and model-based graders
  4. Retrieval and citation metrics
  5. Tool-use and trajectory evaluation
  6. Multi-agent system evaluation
  7. Hallucination, contradiction and abstention testing
  8. Red-team and adversarial test suites
  9. MLflow experiment and artifact lineage
  10. Online experiments and business-impact attribution

Labs Build the Aegis evaluation harness; run a controlled prompt/model/retrieval comparison.

Design review Release or reject a candidate using a documented quality gate.

M9 Security, Safety and Financial-Services Governance Ch 81-90
  1. Threat modelling an enterprise AI platform
  2. Prompt injection and indirect injection
  3. Data exfiltration and retrieval poisoning
  4. Identity, entitlements and row-level controls
  5. Secrets and service-to-service authentication
  6. Tenant, portfolio and strategy isolation
  7. Data-loss prevention and sensitive-output handling
  8. Model risk, validation and inventory evidence
  9. Audit trails, records and defensible reconstruction
  10. Kill switches, approvals and segregation of duties

Labs Attack and defend a retrieval agent; implement entitlement-aware tools and immutable audit events.

Design review Security and model-risk go/no-go review.

M10 Distributed Platform Engineering Ch 91-100
  1. Service boundaries and platform APIs
  2. Synchronous, asynchronous and event-driven execution
  3. Queues, backpressure and admission control
  4. Distributed state, consistency and checkpoints
  5. Caching and semantic-cache hazards
  6. Rate limits, concurrency and workload isolation
  7. Resilience patterns and graceful degradation
  8. Capacity planning for tokens, GPU and vector search
  9. Multi-region design and disaster recovery
  10. Performance and cost engineering

Labs Convert the prototype into distributed services; conduct load, saturation and recovery tests.

Design review Capacity and resilience review against explicit SLOs.

M11 Containers, Kubernetes and GitOps Delivery Ch 101-110
  1. Production container construction
  2. Software supply-chain and image security
  3. Kubernetes primitives for AI workloads
  4. Helm chart structure and configuration strategy
  5. Secrets, workload identity and network policy
  6. Autoscaling and resource governance
  7. GPU scheduling and heterogeneous inference
  8. Flux-based GitOps environments and promotion
  9. Progressive delivery, rollback and migration
  10. Platform developer experience and reusable templates

Labs Package Aegis with Helm; promote it through dev, test and production-like clusters using Flux.

Design review Production-readiness and rollback demonstration.

M12 Observability, Operations and Technical Leadership Ch 111-120
  1. OpenTelemetry traces across agent trajectories
  2. Metrics for models, retrieval, tools and workflows
  3. Structured logging and sensitive-data controls
  4. SLOs, error budgets and alert design
  5. Cost attribution by team, workflow and portfolio
  6. Incident response for AI-specific failures
  7. Drift, degradation and change detection
  8. Platform roadmap and internal product management
  9. Architecture influence and executive communication
  10. Capstone defence and operating-model transition

Labs Build operational dashboards and alerts; run a corrupted-index/failed-model incident game day.

Design review Executive launch decision and 12-month platform roadmap.

Capstone

One scenario, under injected failure

A portfolio manager requests a time-bounded investment review of a listed issuer. The system assembles authorised evidence, calculates portfolio exposure, evaluates a downside scenario, identifies contradictions, drafts a cited investment memo and routes the result for human approval. It must recover from partial service failure without losing auditability.

Failure injections

  • Unavailable primary model
  • Poisoned retrieved document
  • Stale market observation
  • Unauthorised portfolio request
  • Malformed tool output
  • Duplicate workflow event
  • Vector-store latency spike
  • Agent budget exhaustion

Assessment rubric

Business value and workflow fit10%
Architecture and platform reuse15%
Retrieval and data integrity15%
Agent and tool engineering15%
Evaluation quality15%
Security and governance10%
Reliability and operations10%
Deployment and developer experience5%
Technical defence5%

Pass standard: 75% overall, with no score below 60% in retrieval and data integrity, security and governance, or reliability and operations.

Lab environment

Minimum local stack

Early modules run under Docker Compose; later modules move to a local Kubernetes cluster, so the deployment work is done against a real orchestrator rather than described.

  • Python 3.12+, uv or Poetry, Ruff, mypy and pytest
  • PostgreSQL with pgvector, or Qdrant
  • S3-compatible object storage
  • MLflow
  • Airflow or Dagster, and dbt
  • OpenTelemetry Collector, Prometheus and Grafana
  • Docker Compose for early modules
  • Local Kubernetes using k3d, kind or k3s
  • Helm and Flux CLI
How it is delivered

Three ways to take AI Platform Engineering for Investment Management

Self-paced is a document-and-media programme with lifetime access - no live sessions. Cohort and enterprise add live instructor-led training. Mentorship is not offered in any tier.
Feature Self-paced Cohort Enterprise
Format Written chapters, video explainers and podcasts Everything in self-paced, plus scheduled live sessions Everything in cohort, delivered privately to your team
Live sessions None Scheduled, instructor-led Scheduled, instructor-led, private
Mentorship Not offered Not offered Not offered
Access Lifetime Lifetime Lifetime for every enrolled seat
Pace Entirely your own Guided schedule with a peer group Agreed with your desk
Tailoring Fixed curriculum Fixed curriculum Sequenced to your markets, systems and governance
Best for Individuals learning around a job Individuals who want structure and deadlines Desks building the same capability together
For individuals

Get the programme guide

The full chapter list, what each module covers, and how the tiers compare - sent to your inbox as a PDF.

For teams

Run this for my desk

Private delivery for your desk, sequenced to your markets and systems. Tell us the team and we will scope it.

Run this for my desk

Build the platform your desk will actually run

Take it as a cohort, or bring it to your engineering team as a private programme.