Good post -- the within-target Spearman values in Fig 4 look like the real result. Question on the Fig 5 calibration, though. In four of the five panels the points seem to fan away from the diagonal as KDapp increases rather than sitting parallel to it, which would suggest a slope below 1 rather than a pure offset. Target E looks like the clearest case: the points span roughly three logs on x and well under one log on y. Have you fit slopes per target, and do they come out near 1? Worth asking because r and rho are both slope-blind -- E gets rho=0.90 either way -- so nothing in the reported statistics would flag it.
If the slope is materially below 1, an intercept-only correction aligns near the anchor and drifts away from it in both directions, which makes the choice of anchor matter a lot. So, relatedly: how were the five calibration points chosen, and have you run the leave-one-out version, recalculating the offset from each candidate anchor and reporting how the 28/32 count moves? The per-point offsets within a target look wide from the panels -- Target D spans something like two logs across its eleven measurements.
Practically: when a partner runs this on a new target, how many paired BLI measurements do they actually need before the offset stabilizes? One, or is that a demonstration convenience and the real answer closer to three to five?
One suggestion, offered in the spirit of it being the same analysis you're already doing: a hierarchical model with a random intercept per target would give you standard errors on each offset rather than point estimates, and would shrink the small-n targets toward the global mean instead of estimating E's offset from five points. Adding a random slope would answer the first question directly. It also handles the partner case cleanly -- "I've measured two binders on a new target, how much should I trust the offset" is a posterior for a new group, with honest uncertainty attached. With five targets the variance components would be wobbly, but that seems worth knowing too.
You are right that this was intended as an illustrative example. We used a single representative anchor per target and fit only the offset; we did not fit target-specific slopes or run a leave-one-out analysis. The goal was to show that this minimal calibration aligned the measurements well enough for AlphaSeq and orthogonal methods to support similar model-training and decision-making tasks. With more paired measurements, estimating both slope and intercept—or using the hierarchical approach you suggest—would provide a more complete model of the relationship.
That said, we view BLI and SPR as valuable orthogonal comparators, not error-free ground truth. For example, for very tight binders (KD<10nM), slow dissociation can be difficult to resolve accurately, so slope calibration may partly fit AlphaSeq to limitations of the orthogonal assay. In this regime, we and others have found that KinExA can provide more reliable affinity estimates [https://doi.org/10.1080/19420862.2023.2291209]. We have also observed substantial disagreement between KinExA and BLI, which makes us cautious about applying a full slope correction when each method may introduce different biases across the dynamic range. The downside of the KinExA is that the throughput is even lower than SPR and BLI.
On the practical question of how many paired measurements are needed, in a recent collaboration we measured multiple antibodies across multiple targets by both AlphaSeq and BLI, including mutational series for a subset. For the intended use, a target-specific calibration based on the parental antibody was sufficient to recover strong agreement. Since AlphaSeq campaigns typically measure thousands to tens of thousands of antibodies per target, even a half a dozen paired measurements like you suggested remains a highly efficient calibration set.
Thanks for the candid reply, Randolph. Two follow-ups:
1. How were the five anchors in Fig 5 chosen -- were they picked without reference to the other paired measurements in that target (the parental antibody fixed ahead of time, say), or did the choice draw on the rest of the points? If the latter, the offset isn't really coming from one measurement, since selecting the anchor needed the others. Same question applies to the collaboration you mention -- knowing the parental was sufficient required measuring the rest and checking, which is exactly what a partner running this prospectively can't do. Is that why the answer lands at half a dozen rather than one?
Either way, a count alone doesn't seem sufficient, and this is where it comes back to the slope. Say your six anchors all sit at 1–10 nM, and you apply that offset to a campaign where most variants are 100 nM to 1 µM. If the panels fan the way they appear to, the residual out there isn't the same number, and six anchors don't rescue that. Does the guidance include spanning the affinity range, or is it just a count? I'd guess a partner's paired measurements tend to cluster at the tight end, since those are the variants worth confirming, but you'd know better than me.
2. On the BLI limitation: the paper you cite scopes the difficulty to the low picomolar to femtomolar range, but the KDapp-vs.-BLI fan in Fig 4 shows up in targets reading out at tens of nanomolar. Target E looks like roughly 9-32 nM of BLI range against about three logs of KDapp. What would account for the compression there? Related framing question -- if BLI's biases are a reason not to fit a slope, wouldn't the same reasoning apply to fitting the per-target intercept? Both are corrections toward the same comparator. And declining to fit a slope isn't neutral either -- it sets it to exactly 1, with no uncertainty attached. A fitted slope with a wide CI is a weaker claim than an assumed one.
Happy to talk through the hierarchical version offline if that'd be useful -- it'd give you standard errors on the offsets and handle the varying-n across targets, and the partner question ("I've measured two, how much should I trust this") falls out as a prediction interval.
The anchors in Figure 5 were selected retrospectively, with reference to the other paired measurements, to represent the apparent target-specific offset. The calibration therefore did not rely on an independent single measurement. To your point, in a prospective setting, a single point could fall away from the typical offset and produce a less reliable calibration. Estimating the offset from multiple paired measurements would reduce sensitivity to outliers, while allowing for a target-specific slope could further improve the fit. A hierarchical intercept-and-slope model, as you suggest, would be a reasonable way to do both while also quantifying the uncertainty.
Regarding the slope, the KinExA paper does not alone explain the apparent compression in Target E, although I would be cautious about drawing a strong conclusion about its magnitude from the limited number of measurements in Target E. The paper was intended only to illustrate that BLI is not error-free ground truth, particularly for tight binders.
I sent you a connection request on LinkedIn to discuss more offline!
Good post -- the within-target Spearman values in Fig 4 look like the real result. Question on the Fig 5 calibration, though. In four of the five panels the points seem to fan away from the diagonal as KDapp increases rather than sitting parallel to it, which would suggest a slope below 1 rather than a pure offset. Target E looks like the clearest case: the points span roughly three logs on x and well under one log on y. Have you fit slopes per target, and do they come out near 1? Worth asking because r and rho are both slope-blind -- E gets rho=0.90 either way -- so nothing in the reported statistics would flag it.
If the slope is materially below 1, an intercept-only correction aligns near the anchor and drifts away from it in both directions, which makes the choice of anchor matter a lot. So, relatedly: how were the five calibration points chosen, and have you run the leave-one-out version, recalculating the offset from each candidate anchor and reporting how the 28/32 count moves? The per-point offsets within a target look wide from the panels -- Target D spans something like two logs across its eleven measurements.
Practically: when a partner runs this on a new target, how many paired BLI measurements do they actually need before the offset stabilizes? One, or is that a demonstration convenience and the real answer closer to three to five?
One suggestion, offered in the spirit of it being the same analysis you're already doing: a hierarchical model with a random intercept per target would give you standard errors on each offset rather than point estimates, and would shrink the small-n targets toward the global mean instead of estimating E's offset from five points. Adding a random slope would answer the first question directly. It also handles the partner case cleanly -- "I've measured two binders on a new target, how much should I trust the offset" is a posterior for a new group, with honest uncertainty attached. With five targets the variance components would be wobbly, but that seems worth knowing too.
Thank you for the thoughtful comment, Adam.
You are right that this was intended as an illustrative example. We used a single representative anchor per target and fit only the offset; we did not fit target-specific slopes or run a leave-one-out analysis. The goal was to show that this minimal calibration aligned the measurements well enough for AlphaSeq and orthogonal methods to support similar model-training and decision-making tasks. With more paired measurements, estimating both slope and intercept—or using the hierarchical approach you suggest—would provide a more complete model of the relationship.
That said, we view BLI and SPR as valuable orthogonal comparators, not error-free ground truth. For example, for very tight binders (KD<10nM), slow dissociation can be difficult to resolve accurately, so slope calibration may partly fit AlphaSeq to limitations of the orthogonal assay. In this regime, we and others have found that KinExA can provide more reliable affinity estimates [https://doi.org/10.1080/19420862.2023.2291209]. We have also observed substantial disagreement between KinExA and BLI, which makes us cautious about applying a full slope correction when each method may introduce different biases across the dynamic range. The downside of the KinExA is that the throughput is even lower than SPR and BLI.
On the practical question of how many paired measurements are needed, in a recent collaboration we measured multiple antibodies across multiple targets by both AlphaSeq and BLI, including mutational series for a subset. For the intended use, a target-specific calibration based on the parental antibody was sufficient to recover strong agreement. Since AlphaSeq campaigns typically measure thousands to tens of thousands of antibodies per target, even a half a dozen paired measurements like you suggested remains a highly efficient calibration set.
Thanks for the candid reply, Randolph. Two follow-ups:
1. How were the five anchors in Fig 5 chosen -- were they picked without reference to the other paired measurements in that target (the parental antibody fixed ahead of time, say), or did the choice draw on the rest of the points? If the latter, the offset isn't really coming from one measurement, since selecting the anchor needed the others. Same question applies to the collaboration you mention -- knowing the parental was sufficient required measuring the rest and checking, which is exactly what a partner running this prospectively can't do. Is that why the answer lands at half a dozen rather than one?
Either way, a count alone doesn't seem sufficient, and this is where it comes back to the slope. Say your six anchors all sit at 1–10 nM, and you apply that offset to a campaign where most variants are 100 nM to 1 µM. If the panels fan the way they appear to, the residual out there isn't the same number, and six anchors don't rescue that. Does the guidance include spanning the affinity range, or is it just a count? I'd guess a partner's paired measurements tend to cluster at the tight end, since those are the variants worth confirming, but you'd know better than me.
2. On the BLI limitation: the paper you cite scopes the difficulty to the low picomolar to femtomolar range, but the KDapp-vs.-BLI fan in Fig 4 shows up in targets reading out at tens of nanomolar. Target E looks like roughly 9-32 nM of BLI range against about three logs of KDapp. What would account for the compression there? Related framing question -- if BLI's biases are a reason not to fit a slope, wouldn't the same reasoning apply to fitting the per-target intercept? Both are corrections toward the same comparator. And declining to fit a slope isn't neutral either -- it sets it to exactly 1, with no uncertainty attached. A fitted slope with a wide CI is a weaker claim than an assumed one.
Happy to talk through the hierarchical version offline if that'd be useful -- it'd give you standard errors on the offsets and handle the varying-n across targets, and the partner question ("I've measured two, how much should I trust this") falls out as a prediction interval.
The anchors in Figure 5 were selected retrospectively, with reference to the other paired measurements, to represent the apparent target-specific offset. The calibration therefore did not rely on an independent single measurement. To your point, in a prospective setting, a single point could fall away from the typical offset and produce a less reliable calibration. Estimating the offset from multiple paired measurements would reduce sensitivity to outliers, while allowing for a target-specific slope could further improve the fit. A hierarchical intercept-and-slope model, as you suggest, would be a reasonable way to do both while also quantifying the uncertainty.
Regarding the slope, the KinExA paper does not alone explain the apparent compression in Target E, although I would be cautious about drawing a strong conclusion about its magnitude from the limited number of measurements in Target E. The paper was intended only to illustrate that BLI is not error-free ground truth, particularly for tight binders.
I sent you a connection request on LinkedIn to discuss more offline!