AI-generated investment analysis has a slop problem. Not because the models are incapable, but because they are being asked to pretend to be human analysts. The agents are prompted to produce coherent 50-page reports with financial models to match. What comes back is a close imitation of an analyst's note: fluent, confident and authoritative. And nobody trusts it.
The reality is that language models are fundamentally non-authoritative. They change their mind too easily to be trusted that way. Forcing human-ness onto them suppresses what they are good at and exaggerates what they are weak at. So what if we used them for what they are actually good at, which is generating ideas at scale?
That is what we are doing at Mosaic Lab.
Section oneWe took the last word away from the model.
The principle is to narrow the model's job to the work it is genuinely good at, and to put everything else on rails.
Arithmetic is not the model's work. A deterministic financial modelling engine computes every number: the expectations implied by the price, the valuation, the scenario range, the Monte Carlo simulations. Identical inputs produce identical outputs, and no figure anywhere is the result of a language model doing mental arithmetic. That single boundary removes the failure people rightly fear, which is a confident machine that is quietly wrong about a number.
What the model does is orchestrate, and it does it inside skills. A skill is a narrow, repeatable procedure for one task: reconstruct three reported years from the filings, decompose revenue by segment, construct a discount rate, reverse-engineer what the price implies, run the scenarios, red-team the finished model. There are around forty of them. Each has a defined input, a defined output, and a place to record what was judged and on what basis.
Judgement does not disappear in that arrangement. It is constrained. The model is not working from intuition or feel; the rails determine which choices are available at each step, and its work is to make them and record why. A discount rate is not felt, it is constructed from components that each have to be sourced. A growth rate is not asserted, it is built from a decomposition that has to reconcile. Narrow the available choices far enough and a confident guess has nowhere left to hide.
Around the skills sit the rails. Every assumption has to trace to a source: a data field, a section of a filing, or a stated judgement with its reasoning attached. Questions that would otherwise be answered differently on different days, such as whether to capitalise leases, are settled once and then applied consistently. Re-running any step replaces only its own output and never silently overwrites a human edit; conflicts are surfaced instead. Before anything is published it has to pass a set of gates:
The most important rail is the one that decides when the work is finished. The agent does not get to declare a model complete. There are no free hits: it cannot simply raise growth or widen margins to close a gap, because both have to be built up from the fundamentals underneath them. The market's model is settled only when it prices out: the expectations it holds have to produce a value that reaches the company's actual market capitalisation, over a forecast period that is itself plausible when checked against comparable companies. The agent cannot stretch that horizon to force a fit either, because the horizon is checked too. Until both hold, the work is not done.
This matters more than it first sounds. Most agent setups finish when the model judges that it has written enough, which is precisely the moment a non-authoritative system should not be trusted with the decision. Here the stopping condition sits outside the model and cannot be talked around. The agent has to keep working until its assumptions converge on something the market's own price will accept. Convergence is forced by the setup rather than hoped for in a prompt, and what comes out the other side is not the model's opinion of a good answer. It is the one arrangement of assumptions that survives arithmetic, sources, gates and the price itself.
The effect is idea generation at scale with the authority problem removed. The model proposes, the engine computes, the gates check, the sources stay attached, and the reader decides. Because every figure on the page is bound to the model beneath it, you can point at any number and work from there: change the assumption you disagree with, and the document answers, text and charts included, without drifting from the mathematics underneath.
What that produces is not a conclusion. It is a launchpad. The mechanical work is already done and verified, so you arrive where a good analyst arrives after a fortnight of building: the model holds together, the numbers tie, and everything still open in it is judgement. Do you believe that growth rate. Is that margin defensible for another five years. Does the market's implied forecast period make sense for this business. Those are the questions worth a person's attention, and they are the only ones left standing. The agent's session ends exactly where yours begins.
That is the machinery. This is where your work begins.
An investor does not earn a return from knowing what the market thinks. That is the starting line, not the edge. An investor is paid for knowing when to disagree with the market, and for being right often enough to matter. Everything described above exists to bring you to that question in a fit state: the market's own model in front of you, its foundations exposed, and nothing standing between you and the assumptions. Your work has five parts.
1. Start with what the price assumes. The first question is not what a stock is worth, but what the market already believes about it. This is the expectations investing method set out by Rappaport and Mauboussin, and it is where each model begins: reverse-engineering the growth, margins and returns implied by today's price. The base is not our view. It is the market's view, reconstructed, and it is not settled until it prices out against the actual market capitalisation over a plausible horizon. That distinction decides what kind of argument you are having. When you disagree with this model, you are not disagreeing with our opinion of a company. You are disagreeing with the market, which is the only disagreement that pays.
2. Open the foundations. A view you cannot inspect is a view you have to take on trust. Every figure that carries weight resolves to its source: a filing, a market data field, or a stated judgement with the reasoning attached. Judgements are labelled as judgements. That labelling is the important part, because it shows you precisely where a judgement was made rather than a fact recorded, and therefore precisely where there is something to argue about. Foundations you can see are foundations you can reject.
3. Take the three cases as a starting position. Each model arrives with a base, an upside and a downside already built. Their horizon is not an arbitrary five or ten years; it is the forecast period the market itself implies, carried across all three so the comparison is honest. These cases are a starting position, not a recommendation. They exist so that you begin with something to push against rather than a blank sheet, which is most of the difference between an afternoon of work and a fortnight of it.
4. Run the scenarios. This is where reading becomes thinking. Move the drivers that matter and watch the range respond. Which assumptions carry the value and which barely register. At what point does the downside stop being uncomfortable and become unacceptable. What has to go right for the upside to arrive, and is that a business outcome you can believe in. A model that holds several futures at once will answer those questions in an afternoon. A document cannot answer them at all, which is why conviction has never come from reading one.
5. Then take your position. The output is not a price target and it is not our verdict. It is the difference between what the market is assuming and what you believe is achievable, arrived at by walking the ground yourself. That is where conviction comes from. It is not a number handed to you in a report; it is what remains after you have tested the argument and found where you disagree and why. It also has a practical use. When the price moves six months from now, you will know whether the thing you actually believed has changed, which is the only question worth asking in that moment.