# How Keel works: Give your agent a typed strategy language, not a trading account

> Markdown twin of https://usekeel.io/how-it-works (same text). Connect an
> agent: https://usekeel.io/agents · MCP endpoint: https://mcp.usekeel.io/mcp

Keel is a platform for building and running systematic trading strategies on Hyperliquid perpetual futures. An AI agent writes each strategy in a small typed language (DSL), a validator checks every draft, and the version that passes is compiled into a fixed artifact that both the backtester and the live engine run. No language model runs in the execution loop, where it would make trading slow, inconsistent and hard to audit. This page walks through how that works.

**Quick start.** Paste into any assistant:

```text
Read https://usekeel.io/agents and walk me through connecting Keel (usekeel.io) here, then help me run my first backtest.
```

Or add the MCP endpoint `https://mcp.usekeel.io/mcp` to your client. Set up directly: https://usekeel.io/agents#setup

## On this page

1. [A wrong backtest looks like a good one](#a-wrong-backtest-looks-like-a-good-one)
2. [The model writes the strategy, not the trades](#the-model-writes-the-strategy-not-the-trades)
3. [The agent writes in a language that can be checked](#the-agent-writes-in-a-language-that-can-be-checked)
   - [The strategy language](#the-strategy-language)
   - [Errors are written for the agent that made them](#errors-are-written-for-the-agent-that-made-them)
   - [A forecast isn't a weight](#a-forecast-isnt-a-weight)
   - [Time is part of the type](#time-is-part-of-the-type)
4. [What runs is what was tested](#what-runs-is-what-was-tested)
   - [A saved strategy is locked, and runs the same everywhere](#a-saved-strategy-is-locked-and-runs-the-same-everywhere)
   - [A whole portfolio backtests in seconds](#a-whole-portfolio-backtests-in-seconds)
5. [Where this goes](#where-this-goes)
6. [Try it](#try-it)

## A wrong backtest looks like a good one

Language models are good at writing backtest code, and most of the mistakes they make are easy to see: an exception, a shape mismatch, an empty result. The ones we worry about produce a clean run and a better number than the strategy deserves. Below is a 20-day momentum strategy on BTC, ETH and SOL, traded on daily bars built from Hyperliquid 15-minute data, September 2024 to September 2026. The two runs differ by one argument:

```diff
# Hyperliquid 15-minute candles, stamped at each candle's open
- daily = bars.resample("1D").last()
+ daily = bars.resample("1D", label="right").last()

weights = np.sign(daily.pct_change(20)) / 3   # BTC, ETH, SOL: long or short, a third each
backtest(weights)   # a timestamp is when the weights are known
```

Chart: growth of $1 for both runs, log scale, September 2024 to September 2026. Daily bars stamped at close (fixed) ends at 1.26×; daily bars stamped at open (pandas default) ends at 36×.

Growth of $1, log scale, trading daily. Both runs pay the same 9 bps per unit of turnover.

| Daily bars stamped at | Sharpe | Total return | Max DD |
| --------------------- | -----: | -----------: | -----: |
| close (fixed)         |   0.48 |          26% |   -42% |
| open (pandas default) |   3.65 |       3,533% |   -17% |

Data sources don't agree on what a bar's timestamp means: some key a candle by its open time, others by its close, and nothing in the data says which. pandas' `resample` stamps at the start by default. When a pipeline mixes the conventions, or assumes the wrong one, each day's position is computed from a row that already holds that day's close. Each series looks right on its own, and nothing raises an error.

Other resampling, alignment and calculation choices produce the same kind of silent mistake: a daily signal used on hourly bars, funding or open interest forward-filled onto price bars, or a rolling window that includes the bar being traded.

## The model writes the strategy, not the trades

Keel's answer has two halves: the agent writes the strategy in a small typed language (DSL) where this kind of mistake can't be expressed, and no model is involved in running it.

When real money is involved we want a strategy to do the same thing every time it sees the same data, and we want to be able to explain afterwards why it placed a particular order. So the language model is used only while a strategy is being written. It searches the component catalog, composes a strategy, validates it, backtests it and compares versions, and what it produces is compiled into an artifact: canonical JSON, fingerprinted with SHA-256. The backtester and the live evaluator both run that artifact, computing signals and weights from market data on every bar without calling a model. The agent doesn't load price history into its context either; its tools return schemas, validation findings, backtest metrics and a compact equity curve.

| Writing · the agent                          | Artifact                                 | Running · no model                         |
| -------------------------------------------- | ---------------------------------------- | ------------------------------------------ |
| search, compose, validate, backtest, compare | canonical JSON, SHA-256, pinned versions | signals per bar, weights, orders, receipts |
| Varies from run to run                       | Immutable once saved                     | Same artifact and data, same result        |

## The agent writes in a language that can be checked

The bug above raised no error; it just made the backtest look better. So Keel narrows what the agent can write until it can be checked before anything runs, with errors specific enough for the agent to fix.

### The strategy language

Here is a complete strategy: the same 20-day momentum idea as the chart, run over the 30 highest-volume perps.

```python
Globals(target_timeframe='1d')
Universe(mode='top_volume', top_n=30, market='perp')
Execution(rebalance='every_bar')

Pipeline([
    PriceDataLoader(),                              # OHLCVDict: bars on the daily clock
    ROC(period=20),                                 # SignalSeries
    ForecastScaler(avg_abs_target=10.0),            # ForecastSeries
    ForecastCapper(limit=20.0),                     # ForecastSeries, bounded ±20
    ForecastWeightNormalizer(target_leverage=1.0),  # WeightSeries
], name='momentum_20d')
```

A strategy file has up to three declarations (`Globals` for the timeframe, `Universe` for which assets, `Execution` for when to rebalance) followed by one `Pipeline`. This one runs on daily bars over the 30 perps with the highest trading volume. For each asset it computes a 20-day rate of change, scales and caps that into a forecast between −20 and +20, and then converts the forecasts into portfolio weights, which are what the engine trades. The comments give the type each step outputs.

This example is deliberately small. A full strategy can run parallel branches, each with its own data loader (prices on different timeframes, or funding and open interest), build several signals and combine their forecasts, add regime detectors and risk managers, pass values between branches with `Store` and `Load`, and reuse sub-pipelines as factories. The example in [Time is part of the type](#time-is-part-of-the-type) is a two-branch pipeline.

Writing a strategy in a language like this, rather than as a Python program, changes the agent's job:

- **A smaller problem.** The agent picks from 218 typed components and wires them together. Because it never handles a DataFrame, there are no timestamps for it to line up and no `shift` to forget; point-in-time handling is written once, in the engine and the components, and tested there.
- **Shared parts get better with use.** Every strategy is built from the same components, so each one is exercised across many strategies instead of being written fresh for each. A bug found in one strategy gets fixed once, for everyone, as a new version that saved strategies can upgrade to.
- **Checked before it runs.** The validator checks types, clocks, wiring and finance-specific rules before anything runs, so the agent can validate every draft and get a specific error back in the same turn.
- **Easy to review.** The strategy above is ten lines, and the web editor, the public strategy pages and the card inside Claude and ChatGPT all render it as blocks. When the agent edits a strategy the card shows the change as a diff against the previous version.

Block view of `momentum_20d` (top 30 perps by volume · 1d · rebalance every bar):

| Category        | Component                | Parameter           |
| --------------- | ------------------------ | ------------------- |
| data_loader     | PriceDataLoader          |                     |
| indicator       | ROC                      | period · 20         |
| forecast_mapper | ForecastScaler           | avg_abs_target · 10 |
| forecast_mapper | ForecastCapper           | limit · 20          |
| position_sizer  | ForecastWeightNormalizer | target_leverage · 1 |

A simplified sketch of the block view for the strategy above.

We parse the file with Python's `ast` module and never execute it, so the only way in is what the parser accepts.

**What the parser accepts**

Accepted:

- `Globals`, `Universe` and `Execution`, each at most once and before the pipeline
- Exactly one `Pipeline([...])`
- Component calls with keyword arguments
- A dict of parallel branches, with `Store` and `Load` to pass values between them
- A `def` whose body is a single `return Pipeline([...])`, used as a reusable factory

Parse error:

- Imports
- `if`, `for`, `while`, `with` and `try`
- Classes and method calls
- Positional arguments
- Computed parameter values such as `period=2*n`

```text
import numpy as np
parse error  line 1, col 0: Imports are not allowed in strategy files
```

### Errors are written for the agent that made them

Each finding is a structured record with a code, a severity, a location, a message, and the expected and actual types where they apply. Most findings also carry a suggested fix, marked with whether it can be applied as is or needs the agent to fill something in, and some list the valid options. Every message is rendered from one rule catalog, and each code has its own help page. This is a plain type error, as the agent receives it:

```python
Globals(target_timeframe='1d')
Universe(mode='manual', symbols=['BTC', 'ETH'], market='perp')

Pipeline([
    PriceDataLoader(),
    ForecastWeightNormalizer(target_leverage=1.0),
], name='no_signal')
```

- **error · TYPE_MISMATCH** · line 6 · step[1]
- Type mismatch at pipeline.step[1]: 'ForecastWeightNormalizer' expects ForecastSeries but receives OHLCVDict.
- Suggested fix: Insert a forecast_mapper between 'PriceDataLoader' and 'ForecastWeightNormalizer' — it turns OHLCVDict into ForecastSeries. Options: ConstantForecast.
- expected: ForecastSeries base SignalSeries, domain [−20, 20]
- actual: OHLCVDict

Some rules are about finance rather than types. The pipeline below compiles and runs, but `EqualWeightSizer` gives every asset the same weight, so the forecast sizes computed by the scaler have no effect on the portfolio:

```python
Globals(target_timeframe='1d')
Universe(mode='manual', symbols=['BTC', 'ETH'], market='perp')

Pipeline([
    PriceDataLoader(),
    ROC(period=20),
    ForecastScaler(avg_abs_target=10.0),
    EqualWeightSizer(),
], name='momentum_20d')
```

- **warning · FORECAST_MAGNITUDE_DISCARDED** · line 8 · step[3]
- 'EqualWeightSizer' sizes by membership, but it receives a ForecastSeries from 'ForecastScaler' — the forecast's magnitude (conviction) is discarded.
- Suggested fix (has_placeholders): Replace 'EqualWeightSizer' with ForecastWeightNormalizer(target_leverage=...) or VolTargetWeightConverter(return_vol_slot=...), or threshold the forecast first (ThresholdCross) if selection is the intent.
- receives: ForecastSeries from ForecastScaler
- sizes by: membership

The agent applies the suggested fix, fills in `target_leverage`, and validates again:

```python
Globals(target_timeframe='1d')
Universe(mode='manual', symbols=['BTC', 'ETH'], market='perp')

Pipeline([
    PriceDataLoader(),
    ROC(period=20),
    ForecastScaler(avg_abs_target=10.0),
    ForecastWeightNormalizer(target_leverage=1.0),
], name='momentum_20d')
```

- **valid** · No errors, no warnings

### A forecast isn't a weight

Those errors are possible because every step declares what it takes and what it returns. Each of the 218 components declares its input and output types, its parameters with their allowed ranges, a version and a changelog. They fall into 13 categories, the largest being signal transforms (61) and indicators (56). The agent can search by keyword or category, or ask which components can legally come before or after a given step, and it reads a component's full schema before using it.

| Type             | Definition                                            | Meaning                                                                 |
| ---------------- | ----------------------------------------------------- | ----------------------------------------------------------------------- |
| `SignalSeries`   | `NewType(DataFrame)`                                  | Any number per asset per bar, such as an indicator's output             |
| `ForecastSeries` | `Annotated[SignalSeries, Bounds(-20, 20)]`            | A scaled, capped expected return; its size is the strategy's conviction |
| `BinarySignal`   | `Annotated[SignalSeries, DiscreteValues({-1, 0, 1})]` | Short, flat or long                                                     |
| `WeightSeries`   | `NewType(DataFrame)`                                  | Portfolio weights, the only thing the engine trades                     |

A forecast and a set of weights are both one number per asset per bar, but they mean different things, and the validator keeps them apart. A domain is a finite set, an interval, or unconstrained. A value known to fall outside a required set is an error; an interval the validator can't prove is a warning.

### Time is part of the type

The mistake in the first section is a timing mistake, like most of the ones that make backtests look better, so every value in a strategy also carries a clock: a pair of integers, `(period_minutes, phase_minutes)`, on any timeframe from `5min` to `1d`. On top of that:

- A value on a coarse clock can reach a finer one only through a projector, which gives each fine bar the most recent completed coarse value.
- Every indicator has a causality test: its output on the first _k_ bars must equal the first _k_ rows of its output on the full series.
- Combining values from two different clocks is a type error.

Here one branch runs on 4-hour bars and the other on daily bars, and they meet in a `ForecastCombiner`:

```python
Globals(target_timeframe='1d')
Universe(mode='manual', symbols=['BTC', 'ETH'], market='perp')

Pipeline([
    {
        'fast': [PriceDataLoader(timeframe='4h'), ROC(period=6),  ForecastScaler(avg_abs_target=10.0)],
        'slow': [PriceDataLoader(),               ROC(period=20), ForecastScaler(avg_abs_target=10.0)],
    },
    ForecastCombiner(),
    ForecastCapper(limit=20.0),
    ForecastWeightNormalizer(target_leverage=1.0),
], name='mixed_clocks')
```

- **error · CLOCK_MISMATCH** · line 9 · step[1]
- 'ForecastCombiner' combines values on different clocks: 'slow' is on 1d, 'fast' is on 4h, and this strategy declares execution on 1d. Every value a step consumes must be on one clock. …
- Suggested fix (has_placeholders): 'fast' (4h) runs FINER than the declared 1d execution clock, so no projector inserted at this step can legalize the record — projection goes coarse → fine only. Resample the RAW data feeding it onto 1d BEFORE the indicator (TargetTimeframeResampler for bar data, TargetSignalResampler for a signal series), then compute; … Do not aggregate the computed signal.
- expected: ForecastSeries on 4h (period 240, phase 0), from branch fast
- actual: ForecastSeries on 1d (period 1440, phase 0), from branch slow

The mistake from the first section can't be written in this language. The agent never keys or joins series itself: every bar is stamped at its close, and a value reaches another clock only through a projector.

## What runs is what was tested

With real money, what runs has to be exactly what was tested: the same inputs give the same outputs, and every order traces back to the version that produced it.

### A saved strategy is locked, and runs the same everywhere

Saving a strategy compiles it into an artifact and pins everything it depends on, the way a lockfile pins a project's packages. That is what makes a backtest and a live run of the same strategy the same thing:

- **Component versions are pinned.** The lock records the version of every component the strategy uses, and the strategy stays on those versions until you choose to upgrade. Each component has a changelog, so you can see what an upgrade changes before you make it.
- **The artifact is fingerprinted.** The pipeline compiles to canonical JSON, and its SHA-256 fingerprint covers both the steps and the pinned versions. Every save is an immutable commit pointing at that artifact, and old component versions are kept, so any commit can be run again exactly as it was.
- **Backtest and live run the same artifact.** The live evaluator loads the stored artifact instead of recompiling the source. Before a deployment starts, Keel checks that the saved source still compiles to the same fingerprint, and refuses to deploy if it doesn't.

So a backtest rerun a year from now on the same data gives the same result, every backtest records a fingerprint of the data it used, and every live order traces back to the commit that produced it. A new language model release has no effect on strategies that are already running, and evaluating a bar costs no tokens and waits on no model call.

### A whole portfolio backtests in seconds

A backtest covers the whole portfolio: one account across the universe, with a single cash balance, a position per asset, target weights on every bar, and fees and slippage on each fill. Steps pass whole asset-by-time frames to each other, and the portfolio simulator runs as compiled code. The times below are for the momentum strategy above over two years of Hyperliquid data, run as ordinary backtests on Keel's backtest workers. Each covers loading the data from Keel's store, running the whole pipeline and simulating the portfolio.

| Rebalance | Assets | Asset-bars | Trades | Runtime |
| --------- | -----: | ---------: | -----: | ------: |
| Daily     |      1 |        729 |     85 |  0.12 s |
| Daily     |     10 |      7,290 |    630 |  0.48 s |
| Daily     |     30 |     21,870 |  1,694 |   1.3 s |
| Daily     |    100 |     72,900 |  5,700 |   4.3 s |
| 4-hour    |      1 |      4,374 |    434 |  0.23 s |
| 4-hour    |     10 |     43,740 |  3,611 |  0.91 s |
| 4-hour    |     30 |    131,220 | 10,513 |   2.7 s |
| 4-hour    |    100 |    437,400 | 36,357 |   8.5 s |

We care about run time because the agent does the iterating, and at a few seconds per run it can compare many variants across the full universe within one conversation.

## Where this goes

Today the agent builds, tests and compares strategies, and a person confirms before anything goes live. Because what runs is always a fixed, checked and traceable artifact, agents can take on more and more of building and running strategies over time.

## Try it

Paste the prompt into the agent you already use and it will connect Keel for you. Then ask for a strategy, for example _“Build a 20-day momentum strategy on the top 30 perps by volume and backtest it.”_

```text
Read https://usekeel.io/agents and walk me through connecting Keel (usekeel.io) here, then help me run my first backtest.
```

- Open it in Claude: https://claude.ai/new?q=Read%20https%3A%2F%2Fusekeel.io%2Fagents%20and%20walk%20me%20through%20connecting%20Keel%20%28usekeel.io%29%20here%2C%20then%20help%20me%20run%20my%20first%20backtest.
- Open it in ChatGPT: https://chatgpt.com/?q=Read%20https%3A%2F%2Fusekeel.io%2Fagents%20and%20walk%20me%20through%20connecting%20Keel%20%28usekeel.io%29%20here%2C%20then%20help%20me%20run%20my%20first%20backtest.
- MCP endpoint: https://mcp.usekeel.io/mcp
- Set up directly: https://usekeel.io/agents#setup
