Blog / Inside Forecite

[07 / 15]

Mar '267 min read

Most of a model is the label

Two names can print the same 1.8% hour and only one of them moved. Deciding what counts as price impact is harder than fitting anything to it, and it is the reason our scoring model gets rebuilt on a weekly clock instead of a quarterly one.

Yin Nguyen Quant Research

Two items crossed within a minute of each other on a Thursday in February. In the hour that followed, both underlyings were up about 1.8%. One was a mega-cap that trades a 1.8% range on a slow lunch. The other was a name whose entire hourly distribution sits inside half a percent, and 1.8% was the loudest thing it had done all quarter.

Write both of those down as "up 1.8%" and you've taught a model the two events had identical impact. You're wrong twice: wrong about the mega-cap, which did nothing, and wrong about the small cap, which did the most interesting thing on the tape that morning.

People ask me about the weekly cadence. It fits in a sentence, so it's the thing that gets asked. What actually decides whether a scoring model is worth a desk's attention is settled long before anything is fitted, in the definition of what it's being fitted toward. Most of a model is the label.

Did it go up is the wrong target

The obvious supervision target for a news model is the sign of the forward return. Headline lands, price goes one way or the other, mark it up or down, count how often you were right. Intuitive, and close to useless, for three reasons people tend to blur into one.

The first is that a sign is a coin flip wearing a lab coat. Over short horizons the unconditional base rate is near even, so a model can look meaningfully skilled while carrying almost no information, and you'll find out when the fee schedule shows up.

The second is that a binary throws away the size of what happened. An item that moved a name a quarter of a sigma and an item that repriced it entirely get the same tick in the same column. Whatever you fit against that label spends its capacity separating small ups from small downs, because there are vastly more of those, when the thing you needed it to learn was which handful of items were large in any direction.

The third matters most for how we've built the product. A desk doesn't use a news score to pick a side. It uses it to decide what to read. The first question is triage: is this an event, or a company clearing its throat with a timestamp attached. That question is symmetric. A guidance cut and a guidance raise are both extremely actionable, and a label scoring them as opposites is measuring the wrong axis.

So the primary label answers "did this matter", and impact has no sign in it.

Excursion, not endpoint

Once the target is magnitude, the next decision is which magnitude, and this is where a lot of careful work goes wrong.

The default is a point-to-point return: price at publication, price at the horizon, take the difference. One subtraction, and it systematically punishes a model for being right early. An item that pushes a name 3% inside twenty minutes and gives all of it back by the close produced a violent, tradeable, entirely real hour. Its endpoint return is zero, so a model supervised on endpoints learns that event was noise, and next time something of the same shape lands it stays quiet.

Realized excursion fixes that. Record how far the name travelled in each direction over the horizon, favourable and adverse, and use the span. Mean reversion stops erasing the evidence, and a slow grind and a spike no longer collapse into one number.

Excursion comes with an honest limitation. A wide envelope tells you loudly that something happened and very little about which way. It's an excellent label for actionability and a weak one for direction, which is why direction isn't derived from the same measurement: signed short and long direction run from −5 to +5, conviction runs 0 to 100, and they answer different questions off different evidence. A single blended sentiment number hides that. Two columns is not a UI preference. It's an admission about what each one is supervised against.

gate
0100
Everything about the label design is in service of one binary decision a reader makes in under a second.actionability, 0–100

The interesting region isn't the tail. It's the band either side of the gate, where a small change in the score decides whether an item appears on somebody's morning screen at all. That band is also where the measurement is weakest, because a marginal event produces a marginal envelope living close to the noise floor. Evaluation has to look hardest exactly where the ground truth is thinnest.

A percentage is not a unit

Back to the two February items. The fix is simple to state: express the excursion in units of the name's own volatility rather than in percent. A 1.8% hour becomes a fraction of a sigma or a multiple of it, and the mega-cap and the small cap become comparable measurements of the same quantity.

Say it out loud and everyone nods. It's still the most common defect I've seen in news-impact work, including work by people who normalise everything else in their lives. Percent feels like a unit because it has a symbol. It isn't one. It's a raw distance with no denominator, and comparing raw distances across names with a twenty-to-one spread in realized volatility produces a model that has quietly learned to identify small caps.

Time of day needs the same treatment. Hourly excursion isn't stationary across a session, and an item printing in the first half hour is measured against a much fatter background than one printing at 13:20. Skip that and your model learns the opening auction.

The horizon clock has to count sessions rather than wall clock. Something crossing seventeen minutes after the close has no market to move for the rest of the evening, and charging it a flat number of elapsed hours makes every Thursday-evening release look inert and every Tuesday-morning one look explosive. The difference is the calendar rather than the news.

One boundary falls straight out of this. Every horizon we supervise against is measured in minutes rather than microseconds, a deliberate statement about what this product is for. If your edge is the first tick, we're not the tool. We're built for the reader deciding, at 07:02, which four of ninety-one rows are worth opening.

Random splits leak the future

Now the evaluation side, where the failure is subtler and worse.

Shuffle your rows, hold out 20%, fit on the rest. That's the default in every tutorial and it's invalid on time series for a reason that has nothing to do with the model. Financial data is cross-sectionally correlated in time. On a heavy earnings morning a hundred items land into the same tape, the same rate expectations, the same risk appetite. Split them at random and most of that morning sits in your training set while a few rows of it sit in your test set. You're grading the model on a morning it has already read.

The estimate that comes back isn't slightly optimistic. It's optimistic in the direction that feels most convincing, because it looks like generalisation across names, which is the thing you were worried about.

Walk-forward is the only defensible answer. Fit on everything up to a date, evaluate on what came after, roll the date, repeat. Every number is out of sample in the one direction that can't be faked. It costs you: less data per fit, noisier estimates because each window is small, and no tuning your way to a better answer, since the tuning is itself part of what needs testing out of sample. Purge the boundary too, or an event whose horizon straddles the split hands you the answer. The estimates come out worse-looking, and they're the only ones that survive contact with a live feed.

Which is why the clock is weekly

I spent four years at a stat-arb shop watching good backtests die. They almost never died from a bug. They died because the relationship they measured stopped being the relationship that was happening.

An in-line guidance print does one thing in a tape pricing a cutting cycle and something else in a tape pricing the opposite. Same text, same sub-scores, different response function. The mapping from what a headline says to what a price does isn't a law of nature. It's a regime-conditional empirical fact, and a model fitted against a year of price behaviour is a statement about that year.

So the Verdict Engine gets rebuilt on a rolling window, and the window rolls every week. Most weeks the candidate looks a great deal like the incumbent, and that's the intended result rather than a wasted cycle. The cadence exists so the gap between a shift in the response function and a model that has seen the shift is measured in days. A promotion path you exercise fifty times a year is one you can trust in the week the regime turns, which is never the week you planned for.

What I want next is a better denominator. Normalising the excursion by the name's own volatility handles the mega-cap against the micro-cap and does nothing about a day when an entire sector moves together and every item in it inherits a wide envelope it didn't earn. The event was the sector, not the filing. Stripping that shared component out before it becomes a label would make a single global gate at 60 mean something closer to the same thing in energy as it does in staples, which it currently does not, and that's the piece of the measurement I'd most like to be arguing about six months from now.