The reality, up front
Compensation, stated plainly
There is no single answer — it depends on the track, and every track is stated honestly rather than dressed up:
- Unpaid research collaboration with co-authorship. You work on an open problem, your contribution is credited, and the work itself (plus the resulting writeup or result) is the compensation.
- Equity partnership. For a deeper, ongoing commitment, equity in the research venture can be discussed directly — terms are individually negotiated, never posted publicly.
- A role that may become paid if and when funding is secured — with no guarantee of timing, and no guarantee that it occurs at all.
The realistic default: today, pre-funding, nearly every slot is unpaid. If you need paid work now, this is not that — and it's better to know that before either of us spends time on a conversation.
No live capital, no revenue
There is currently no live capital deployed and no revenue. There is no near-term funded-employment path independent of a successful raise. This is more decision-relevant than any GPU spec below, which is why it's stated here rather than buried in an FAQ.
A day in the life — the environment
This part is the genuinely engaging reason to work here, and it's real — but every claim below is checked against what's actually wired today, not what's designed on paper.
- You launch from a single frozen config — one source of truth for a run, not a pile of ad-hoc notebooks and hand-edited scripts.
- Automated run-health guardrails abort pathological runs early. You won't waste a weekend on a run that was dead in the first hour. (Guardrails are implemented for the core failure modes; broader coverage is still expanding — designed for more than is live today.)
- Every run produces a deterministic evaluation dossier on held-out data before it's taken seriously.
- An agent advances — sim → paper → canary → full — only by clearing a qualitative set of promotion gates, not by looking good on one chart. (Sim and paper stages are live; canary/full-deployment stages are designed, not yet built — there is no live capital today, see above.)
- You work inside a reproducible pipeline: same seed, same code, same data → the same result, byte for byte. Determinism is a hard requirement, not an aspiration.
In short: a rigorous, guard-railed research pipeline — not a pile of notebooks — but still a one-person infrastructure, so "designed" and "built" are not always the same thing, and this page says which is which.
How your experiment would be run
Every experiment starts from the same template, adapted from the one used for every published result on this site:
- Fixed / independent / dependent variables stated explicitly before the run starts, not reverse-engineered from the result afterward.
- Multiple random seeds where the question requires statistical confidence, not a single lucky run.
- Pre-declared statistical tests (confidence intervals, effect sizes) rather than eyeballing a chart.
- Explicit success criteria written down before you start — including what a negative result looks like, and a commitment to report it as a negative result rather than quietly shelving it.
The hard constraints you'd be fighting
Beating these constraints — not working around them — is the actual job:
- Low signal-to-noise ratio. Financial time series are close to the hardest domain in applied RL for exactly this reason.
- Roughly a decade of daily history per stock. Not the scale of a modern deep-learning dataset — every method has to work with limited data.
- No order-book, fundamental, or alternative data in scope today — price/volume-derived signals only.
- Survivorship bias in the underlying universe is a known, unresolved limitation, not something already solved.
- Single consumer GPU, one workstation, no cloud. Everything has to run efficiently on hardware you could buy yourself.
What you'd own
These are real, currently-open problems — not busywork. Full detail, with the reasoning behind each, lives on Open Research Problems.
- Regime-aware meta-learning — agents that detect and adapt to regime shifts instead of degrading silently.
- Continuous target-weight action spaces — moving beyond discrete position sizing toward continuous portfolio-weight control.
- Offline RL (CQL / IQL) — learning from logged data to reduce reliance on simulation-to-reality transfer.
- Attention / transformer cross-sectional selection — architectures that reason jointly across the stock universe, not one asset at a time.
- Multi-objective / Pareto RL — treating return and risk as a genuine trade-off frontier instead of a single scalar reward.
Role scopes & logistics
- Research collaborator: owns one open problem end-to-end — hypothesis, experiment, writeup.
- ML / RL engineer: works on the training loop, the walk-forward evaluation harness, and pipeline reliability (no internal module or constant names shared outside the codebase — this page describes scope, not implementation).
- Stack in practice: Python, a JAX-based training loop, PyTorch for parts of the pipeline, a walk-forward evaluation harness, Parquet/SQLite for data.
- Remote, async-friendly. No fixed office, no fixed time zone requirement — overlap for regular sync is appreciated, not mandated.
Express interest
This is a lightweight qualifying form, not a raw inbox — it helps focus a reply on your actual question. It opens your own email client with the message pre-filled; nothing is stored on any RLAlphaLabs server. See the privacy notice for details.