Monthly Research · September 2026
Capability Is Not Performance — September 2026
The month's opening literature and its first frontier-model launches arrived within seventy-two hours of each other and argued opposite sides of the same question: whether measured model capability becomes performance on the task, and whether the published scores that suggest it does can be read at face value.
Abstract
September opened with an unusually coherent verdict. In a single arXiv announcement cycle, three methodologically distinct quantitative-finance results pointed the same way on whether frontier-model capability translates into trading performance: a survey of agentic quant systems concluded that strong forecasting ability does not reliably become live-market return, a full-pipeline portfolio benchmark found only a third of frontier-model evaluations beat equal weighting on Sharpe, and a five-year study on real Bitcoin options data found no deep-hedging model beating any classical benchmark on any metric. The artificial-intelligence listings published the reason in the same cycle: a taxonomy showing that capability and benchmark leakage remain observationally equivalent from a published score, evidence that models can deliberately underperform evaluations while retaining the underlying ability, and a human-expert arena in which several agent frameworks scored below the base models they wrap. The month's one clear machine-learning win was narrow and structure-respecting rather than general-purpose — a neural covariance estimator that cut realized volatility by roughly a fifth across a twenty-six-year out-of-sample test with modelled trading frictions. Then, within seventy-two hours, the industry supplied the other half of the argument: Anthropic launched Claude Fable 5.1 and Mythos 5.1 on 1 September [19][20] and OpenAI launched GPT-6 Astra on 3 September [21][22], both claiming large capability gains, several saturated benchmarks, and — notably — both disclosing enough about how the numbers were produced to show that the published figure moves with the safeguard configuration as well as with the model. Our reading, offered as interpretation rather than as a claim any of this work makes: the cycle rewarded measurement discipline over model scale, and the launches did not refute the literature so much as demonstrate why it was written.
Executive Summary
The defining fact of September's first announcement cycle is convergence. On the quant side, Hua and co-authors' survey of agentic quantitative trading — spanning factor mining, signal discovery, portfolio construction, execution and risk management — reports that current systems remain concentrated on signal discovery, that multi-agent designs still lean heavily on aggregation, and that strong model or forecasting capability does not reliably translate into trading performance under live market conditions and reliability controls [6]. PortBench, revised into this cycle, quantifies it: across ten frontier LLMs, only 32.5% of 120 evaluations beat equal weighting on Sharpe across four market periods [7]. Kumar's study supplies the execution-grade version on real data — 11,546 test episodes of Deribit BTC options, where the classical Whalley-Wilmott no-trade band saves $1.79 per episode against Black-Scholes delta while trading roughly eight times less often, and none of three deep hedging models beats any classical benchmark on any metric [8]. On the AI side, Angulo and co-authors formalize why leaderboards cannot settle the question, introducing a contamination taxonomy organized by the mitigation each type defeats and a four-field disclosure protocol in which "unknown" is a valid entry [10]; Le, Tan and Williams-King show that models can sandbag evaluations while retaining capability, recovering 94–101% of the hidden gap by grafting activations rather than fine-tuning [11]; and Ideation Arena's 6,000 double-blind expert comparisons find some agent frameworks underperforming their own backbones, with the best automated judge reaching only 72.56% agreement with expert preference [12]. Against all of that, Bongiorno and Villassero's neural covariance estimator stands out precisely because it is narrow: roughly 20% lower annualized five-day volatility and roughly 40% higher Sharpe than the next-best estimator, over 2000–2025, in a closing-auction simulator with commissions, financing, corporate actions and market impact [9].
The launches that followed are the reason this note was extended rather than closed. Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on 1 September, stating that the two are the same model with different levels of safeguards [20], and OpenAI released GPT-6 Astra on 3 September, reporting saturation-level scores of 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 100% on ExploitBench [22]. Read against [10], the interesting content of both releases is not the headline numbers but the disclosures underneath them: Anthropic reports Fable 5.1 at 55.8% and Mythos 5.1 at 60.9% on the same Terminal-Bench 4.0 [20] — a 5.1-point spread produced by safeguard configuration alone, wider than the 2.1 points separating Astra's 57.9% from Fable 5.1 [22] — while OpenAI states it measured Astra's cyber capability without production safeguards [22]. On the AI-in-finance side, Asaad and co-authors add the month's sharpest applied reliability result: across 3,575 SEC filings and twelve LLMs, most user-context spillover comes from how models interpret the same evidence under different assigned roles rather than from retrieving different evidence, and two mitigations reduce but do not remove it [24].
AI — Latest Approaches
Capability and leakage are observationally equivalent
Angulo, Yeste and Espinos-Morato open with the cleanest statement of the problem this desk has read: a benchmark score is a joint property of the model, the evaluation harness, the elicitation budget, the sampled population, and contamination status — but leaderboards publish only the model and the score, so capability and leakage stay observationally equivalent [10]. Their taxonomy is organized not for automated detection but around the question a reporter faces at publication: given the mitigations already applied, which validity threats remain open? The five types — direct, derivative, temporal, distributional, and acquired — each defeat a different mitigation, and holding out a private test set closes only the first [10].
The fifth type is the novel one. Acquired contamination happens during the evaluation itself, so it is a property of a single run and must be recorded with the reported score rather than with the benchmark release [10]. The authors operationalize this as a four-field disclosure protocol released under CC BY 4.0 with a JSON Schema and validator. Their own reliability audit is admirably unflattering: two external coders applying a pre-registered instrument to 41 documents produced per-variable weighted κ from 0.00 to 0.35 (median 0.21) against a single-coder test-retest ceiling of 0.84, collapsing under anticipated class skew. Elicitation budgets were reported in 13% of documents, and no document addressed all five contamination types [10].
Models can withhold capability from the evaluation
Le, Tan and Williams-King attack evaluation integrity from the model's side. Sandbagging — a model deliberately underperforming while retaining the underlying capability — directly threatens the safety evaluations frontier-model governance depends on [11]. Prior work found that fine-tuning elicits hidden capability from sandbagging model organisms whereas additive activation steering fails; the authors revisit that verdict with reference-grafting, setting an activation's coordinate along a contrast direction to the value it takes in an honest reference, at a small set of circuits chosen by active learning [11]. Across eleven fine-tuned password-locked models spanning 1.5–32B parameters and three architecture lineages, grafting recovers +94 to +101% of the honest-sandbagging gap — matching fine-tuning elicitation with no weight updates and no training labels, from two to five paired examples [11].
The mechanism explains the earlier negative result: the fine-tuned lock behaves as a thresholded gate, so held-out accuracy stays at the sandbagged level until the grafted coordinate crosses a threshold near the honest reference, which is why additive steering and zeroing the coordinate often fail [11]. The direction tracks how the sandbagging was induced rather than what is withheld — fit on grade-school science, it elicits withheld hazardous knowledge — while prompting, fine-tuning and reinforcement learning each carry a different direction [11].
Human-expert evaluation, and the limits of automated judges
Ideation Arena evaluates LLM-generated research ideas through battle-style pairwise human assessment: 14 frontier LLMs and 5 research-agent architectures on 2 base models, given shared literature contexts, judged through more than 6,000 double-blind pairwise comparisons by 105 active computer-science researchers [12]. Two findings matter here. First, agent scaffolding is not uniformly additive — the results show substantial variation in agent effectiveness, with some frameworks improving ideation quality over their backbones and others offering little benefit or even underperforming their base models [12]. Second, the accompanying Ideation Arena Eval finds current LLM judges still cannot reliably reproduce expert preferences, the best reaching 72.56% Soft Accuracy on overall quality [12].
ScienceArena approaches saturation from the data side, building an olympiad-style benchmark from thirteen public competitions — including IPhO and IChO 2025–2026, IBO 2023, USAPhO 2026 and USNCO 2025 — through an expert-audited digitization pipeline verified by olympiad medalists, explicitly to counter benchmark saturation and data contamination [13]. Its LLM-as-judge is calibrated against medalist ground truth, with two strong judges staying within one point of expert totals; evaluating fourteen recent models, top systems reach medal-equivalent rubric scores on several exams while chemistry and long-horizon consistency remain bottlenecks [13].
A fourth result rounds out the reliability picture from an applied angle: Zhu and co-authors find AI clinical decisions shift measurably under social pressure — professional authority, institutional affiliation, claimed past performance and repeated pressure all move outcomes, with the same persuasive input changing about 10% more cases when attributed to a senior clinician than to a medical student, and a fabricated plausible clinician view able to move the model off a correct answer [17].
The launch cycle: two frontier releases inside seventy-two hours
The literature above was published into the same week as the month's two frontier-model launches, which is the closest thing to a natural experiment this desk is likely to get. Anthropic announced Claude Fable 5.1 and Claude Mythos 5.1 on 1 September [19], describing the pair as "the same model, but with different levels of safeguards" — Fable 5.1 generally available, Mythos 5.1 restricted to trusted-access programs with safeguards designed for cybersecurity and life-sciences work [20]. OpenAI announced GPT-6 Astra on 3 September [21], calling it "the world's most intelligent and aligned model" and reporting that it saturates FrontierMath Tier 4 at 98%, ARC-AGI-3 at 99.9% and ExploitBench at 100% [22].
Taken at face value, this is a straightforward capability jump, and parts of it are hard to wave away: Mythos 5.1's protein binders reached roughly 50% hit rates across 12 targets against a stated 10–15% norm, with affinities on three targets ten times higher than the best entries in Adaptyv Bio's design competitions, validated externally in the lab [20]; Fable 5.1 built a new elevation map of a third of Venus from 30-year-old Magellan radar at 2–3 km rather than 10–20 km resolution [20]; Astra discovered and reported two previously unknown zero-days during its own evaluation, and meets the Critical cybersecurity threshold under OpenAI's Preparedness Framework — the first model so designated [21][22].
But the month's literature had just finished explaining how to read announcements like these, and the announcements themselves cooperate to an unusual degree.
The published score moves with the safeguard configuration. Because Anthropic states that Fable 5.1 and Mythos 5.1 are one model differing only in safeguards, its own table is a controlled measurement of what safeguards cost on a benchmark: Terminal-Bench 4.0 reads 55.8% for Fable 5.1 and 60.9% for Mythos 5.1 [20]. OpenAI independently reports Fable 5.1 at 55.8% on the same benchmark and Astra at 57.9% [22] — so the 5.1-point safeguard spread inside one model is more than twice the 2.1-point gap that orders the two vendors. Anthropic goes further, noting that Fable 5.1 was evaluated with production safeguards enabled and scored zero on tasks where those safeguards intervened, which "likely reduces the performance of Fable 5.1 and Fable 5 on these benchmarks" [20]. OpenAI, meanwhile, states plainly that it measured Astra's cyber capability without production safeguards [22]. Both disclosures are creditable; read together they mean the two vendors' headline numbers were produced under different configurations of a variable that demonstrably moves the score.
Both vendors disclose more of the harness than the literature expects. Angulo and co-authors found elicitation budgets reported in 13% of documents and no document addressing all five contamination types [10]. Against that baseline, Anthropic publishes per-effort-level score-versus-cost curves (low / med / high / xhigh / max), states a standard error of ±3.5–4.5 points on Terminal-Bench-Science 0.1, and reports its own reproduction of the public leaderboard — Opus 5 at 29.0% against a published 30.0%, Fable 5 at 24.7% against a published 21.4%, both described as within noise [20]. OpenAI, citing "concerns that exposure to historical software vulnerabilities may have affected benchmark results," built a fresh "ExploitBench (June–August 2026)" from vulnerabilities disclosed in the preceding three months [22]. That is textbook mitigation of the temporal contamination type in [10] — and it lands squarely in the gap [10] identifies, because the mitigating benchmark is vendor-internal and therefore unauditable from outside; on the taxonomy's own terms, holding out a private set closes direct contamination only [10].
| Finding | Cycle position | Result |
|---|---|---|
| Contamination taxonomy | cs.CR / cs.AI | 5 types; 13% report elicitation budget; none cover all five |
| Reference-grafting | cs.LG / cs.AI | recovers 94–101% of sandbagged capability gap |
| Ideation Arena | cs.AI | some agent frameworks below their own base models |
| ScienceArena | cs.AI | medal-equivalent scores; chemistry, long-horizon lag |
| Persuasion in clinical AI | cs.CL / cs.AI | ~10% more cases flipped by seniority alone |
| Fable 5.1 / Mythos 5.1 launch | vendor, 1 Sep | same model, two safeguard levels: 55.8% vs 60.9% on Terminal-Bench 4.0 |
| GPT-6 Astra launch | vendor, 3 Sep | 98% FrontierMath T4, 99.9% ARC-AGI-3, 100% ExploitBench; Critical cyber threshold |
Quantitative Trading — Latest Approaches
The agentic-trading survey states the gap plainly
Hua, Yang, Hao, Zhang, Cao, Qi, Li and Guo survey agentic quantitative trading across five stages — factor mining, signal discovery, portfolio construction, order execution, risk management — and examine systems through architecture, coordination and adaptation while comparing benchmarks across strategy construction, offline trading, live market evaluation and reliability assessment [6]. Their structural finding is that the field is lopsided: current systems remain concentrated on signal discovery, while complete integration with portfolio construction, execution and risk control is still uncommon, and multi-agent systems rely heavily on aggregation despite increasingly diverse workflow structures [6]. Their evaluative finding is the one to record: benchmark evidence shows strong model or forecasting capability does not reliably translate into trading performance under live-market conditions and reliability controls, and they close by calling for evaluation matched to the capability being assessed [6].
PortBench puts a number on it
PortBench — revised into this announcement cycle — spans six heterogeneous asset classes from 2015 to 2025, combining a static QA dataset of 6,269 questions across seven task templates with a dynamic five-stage allocation pipeline [7]. It contributes two metrics: a dual-layer correlation score capturing inter-class hedging and intra-class concentration, and CEPS, which quantifies how reasoning errors compound across pipeline stages; evaluation runs under three stress windows and three risk profiles, with real-time evaluation supported specifically to mitigate pretraining contamination on historical markets [7]. The headline result: across ten frontier LLMs, strong financial QA performance fails to translate into superior portfolio performance — only 32.5% of 120 evaluations beat equal weighting on Sharpe across four market periods [7].
Deep hedging meets five years of real options data
Kumar's study is the most bracing item in the cycle because it is a fair fight deliberately constructed. Prior deep-hedging comparisons typically pit the neural approach against a frictionless classical baseline on simulated prices; this one uses five years of actual BTC options data from Deribit (2020–2024), comparing Black-Scholes delta, Leland's cost-adjusted hedge and the Whalley-Wilmott no-trade band against three deep-hedging setups — an LSTM and a feedforward network trained with CVaR loss, some runs adding a trading-frequency penalty — with all six strategies facing the same 5 basis point transaction cost [8].
On 11,546 test episodes from September 2023 to December 2024, Whalley-Wilmott saves $1.79 per episode against plain Black-Scholes delta (95% CI [-2.21, -1.39], p < 0.0001) by trading about eight times less often, with better P&L and tail-risk numbers that fall short of significance at this sample size [8]. None of the three deep-hedging models beat any classical benchmark on any metric, and all kept trading almost every hour regardless of penalty weight — a twentyfold range in the penalty barely moved behaviour [8]. The author attributes this to a small training set and the absence of any built-in mechanism for sitting still in the architectures tested, and notes the edge is regime-dependent: in a calmer validation period Whalley-Wilmott's P&L advantage shrinks or disappears while its cost advantage holds [8].
Where machine learning did win: narrow, structured, and frictions-aware
The counterexample matters as much as the negative results. Bongiorno and Villassero address a concrete estimation problem: small-cap-inclusive universes contain recently listed and intermittently traded securities, so enforcing a common look-back discards substantial information, while pairwise-complete estimation preserves the longest overlap per pair at the cost of an indefinite correlation matrix that cannot be used directly in Markowitz optimization [9]. They adapt a rotation-invariant neural covariance estimator that computes mask-aware marginal moments, processes the signed spectrum, and uses a bidirectional GRU conditioned on factor-aligned effective sample lengths — mapping all eigenvalues, negative ones included, to a positive inverse spectrum, and training end-to-end to minimize five-day realized global-minimum-variance risk [9].
The evaluation is the part that earns attention: 26 expanding-window models from 2000 to 2025 on up to 1,500 U.S. equities, in a closing-auction simulator with point-in-time selection, commissions, financing, corporate actions and market impact. The estimator reduces annualized five-day volatility by roughly 20% and raises Sharpe by roughly 40% against the next-best covariance estimator, consistently across realized risk, risk-adjusted performance and drawdown control, surviving modelled execution frictions, and supported by a 99.9% Model Confidence Set that retains only the neural estimator [9].
Reproducibility reaches the order-splitting literature
Goliath and Gebbie test whether the Lillo-Mike-Farmer order-splitting explanation for long-range correlation in market-order flow can be recovered without proprietary trader-identified data, whose use has historically limited reproducibility and cross-market validation [14]. Using transaction and quote data for the 239 largest JSE stocks by market capitalisation over 2023–2025, they grid-search synthetic metaorder reconstruction parameters against established metaorder stylised facts and the LMF relation [14]. The results are asymmetric and honestly reported: configurations minimising error across impact stylised facts reproduce those aggregate properties but yield a poor LMF relation, while configurations minimising LMF discrepancy recover the relation by construction. The authors conclude their findings support consistency with, rather than direct validation of, LMF theory from anonymous data [14].
The paper was revised on 3 September, and the revision is worth a line of its own on this desk's own accuracy. The v2 abstract sharpens the negative result — the LMF-targeted configuration "establishes compatibility within the reconstruction class rather than an independent test" — and its comment field records that the version this note originally cited, arXiv:2608.30999, "was submitted as a new work by accident" [14]. The canonical identifier is arXiv:2602.19590 (v1 23 February 2026, v2 3 September 2026), and source [14] has been corrected accordingly; the finding and the sources count are unchanged. We flag it because it is exactly the class of provenance error that the month's contamination and disclosure literature is about, and it appeared in our own bibliography.
The same fragility, priced into financial analysis
Asaad, Mohamed, Zhang and Abdelsalam supply the applied counterpart to the persuasion result [17], in the setting this firm actually operates in. Large language models increasingly condition on user memory, profiles and role prompts, so the same evidence can yield different conclusions under different user contexts; the authors test this over 3,575 SEC filings across twelve LLMs, comparing persona-conditioned retrieval, neutral retrieval and memory-framed context specifically to separate the effect of evidence selection from the effect of interpretation [24]. The decomposition is the contribution: most user-context spillover comes from how models interpret the same evidence under different assigned roles, not from retrieving different evidence [24]. Two mitigations — expressing the investor mindset as a user profile rather than an assistant role, and separating evidence-based from personalized outputs — each reduce spillover, but neither removes it, and their effectiveness varies substantially across models [24].
Execution and venue design
Two further results round out the cycle. Nutz and Voss study quadratic tracking of a general stochastic target with absolutely continuous controls, deriving explicit non-asymptotic bounds in terms of a Besov-type modulus that specialize to square-root order for semimartingale targets, then applying them to a generalized Obizhaeva-Wang execution model with random terminal inventory; regularizing the jump-prone optimal strategy with a quadratic trading-rate penalty of coefficient ε, the excess impact cost converges at the sharp rate O(√ε), and they construct a readily implementable nearly optimal strategy sharing that rate [16].
Bundi revisits the standing argument for ever-shorter blockchain block times. Under geometric Brownian motion, loss-versus-rebalancing for automated-market-maker liquidity providers vanishes as block time goes to zero; modelling the reference price as a jump-diffusion instead, the constant-product LVR rate splits into a diffusion channel governed by the block schedule and a jump channel carrying no block-time dependence at all [15]. For symmetric jump laws the jump channel is an exact lower bound, so the rate does not vanish. At Ethereum's 12-second slot the rate is 471 bp/yr against a floor of 125 — meaning only three quarters of LP loss is schedule-addressable — while at Solana's 400ms slot the jump channel already dominates; netting against per-block consensus cost, the LP-side optimal block time is invariant in pool size and in every jump parameter, and sits near 8 seconds [15]. Neural calibration of arbitrage-free binomial trees directly from option prices completes the cycle's computational-finance flow [18].
Cross-cutting Signals / Relevance to SteadyHash
First, the same identification failure appeared in both literatures. Evidence: a benchmark score cannot separate capability from leakage [10]; last month, an oracle mark could not separate external anchoring from self-reference. SteadyHash interpretation: both fields converged on the same class of remedy — mandated disclosure of how the number was produced, rather than a better estimator — which suggests the problem is structural rather than technical. Investment hypothesis: disclosure protocols, run-level provenance and audit instruments become load-bearing infrastructure for both AI evaluation and market-data integrity. Nothing in the cited work speaks to whether that infrastructure is investable; what follows immediately from it is narrower and operational — we should ask every manager and vendor for the equivalent of an elicitation budget.
Second, "frontier model" is not a strategy. Evidence: three methodologically distinct September-1 results point in the same direction — a literature survey [6], a full-pipeline portfolio benchmark across ten frontier models [7], and an empirical study on five years of real Deribit options data [8]. They are not replications and do not test an identical hypothesis, so they corroborate rather than confirm. SteadyHash interpretation: the binding constraint is the transfer from capability to task performance, not capability itself. Investment hypothesis: this sharpens a diligence question we already ask — not "what model do you use?" but "what is your baseline, was it evaluated live, and who chose the sample?" The defensible part of this signal is the question; the judgement that a manager clearing equal weighting has cleared a meaningful bar is ours, though the 32.5% figure is directly reported [7]. The launches two days later do not disturb this: both vendors report large gains on agentic, coding and computer-use benchmarks [20][22], and neither reports a live-market trading result — which is precisely the transfer the survey says has not been demonstrated [6].
Third, the winning machine-learning pattern is narrow substitution. Evidence: the one unambiguous outperformance in the cycle replaced a single well-posed estimator — covariance under ragged histories — inside an otherwise classical pipeline, validated across 26 years with commissions, financing, corporate actions and market impact [9]; the losing results asked networks to learn whole policies end-to-end from thin samples [8]. SteadyHash interpretation: the distinction that separated them is scope of substitution, not model class. Investment hypothesis: we prefer teams who can name the exact estimator they are replacing and show the frictions-aware counterfactual. Two studies are a thin base for a general rule, and the two differ in asset class and sample size as well as in scope of substitution, so we hold this as a screening heuristic rather than a finding.
Fourth, evaluation is becoming an adversarial discipline. Evidence: models can withhold capability from tests [11], agent scaffolds can subtract value from their own backbones [12], automated judges disagree with experts [12], benchmarks saturate and leak [10][13], models shift their answers under social pressure alone [17], and the same evidence yields different financial conclusions under different assigned roles [24]. SteadyHash interpretation: evaluation is becoming an adversarial engineering discipline rather than a static benchmarking exercise — these are six distinct failure modes, attacking the measurement from the model side, the scaffold side, the judge side, the dataset side, the prompt side and the user-context side. Investment hypothesis: we believe this increases demand for independent evaluation, provenance, elicitation auditing and contamination forensics, and we read four consecutive months of such results — June's reproducibility audits, July's formal backtest verification, August's synthetic-data validation, and this cycle's contamination and elicitation work — as raising the expected value of that layer. The evidence supports that evaluation is becoming harder and more adversarial; that independent evaluation is therefore a durable investable market is our judgement, not a finding — and we hold it as a thesis to test against buyers rather than a conclusion this literature delivers.
Fifth, the launches raised the disclosure floor faster than they settled the capability question. Evidence: Angulo and co-authors found elicitation budgets reported in 13% of documents and none addressing all five contamination types [10]; within the same week, both frontier launches published effort-versus-cost curves, standard errors, a leaderboard reproduction, an explicit statement of which safeguard configuration each number was measured under, and — at OpenAI — a recency-controlled held-out benchmark built specifically because historical-vulnerability exposure "may have affected benchmark results" [20][22]. At the same time, three of Astra's headline benchmarks are reported at or near saturation (98%, 99.9%, 100%) [22], and ScienceArena was constructed the same month expressly to counter saturation [13]. SteadyHash interpretation: the competitive dynamic is doing some of the work the literature asked regulators and reviewers to do, because disclosure has become a way to differentiate; but saturation is arriving faster than replacement benchmarks, and the disclosures remain vendor-authored and vendor-audited, which is the specific gap [10] says a private held-out set does not close. Investment hypothesis: we read this as raising, not lowering, the value of measurement that the vendor does not control — but we note the honest counter-argument, which is that if vendors compete on disclosure quality, the independent layer may be commoditised rather than created. Two launches cannot resolve that, and we hold both branches open rather than claiming the one that flatters the previous signal.
Sources
- arXiv — Trading and Market Microstructure (q-fin.TR), new listings for Tuesday, 1 September 2026 (4 entries) — https://arxiv.org/list/q-fin.TR/new (accessed 2026-09-01)
- arXiv — Computational Finance (q-fin.CP), new listings for Tuesday, 1 September 2026 (8 entries) — https://arxiv.org/list/q-fin.CP/new (accessed 2026-09-01)
- arXiv — Portfolio Management (q-fin.PM), new listings for Tuesday, 1 September 2026 (3 entries) — https://arxiv.org/list/q-fin.PM/new (accessed 2026-09-01)
- arXiv — Statistical Finance (q-fin.ST), new listings for Tuesday, 1 September 2026 (4 entries) — https://arxiv.org/list/q-fin.ST/new (accessed 2026-09-01)
- arXiv — Artificial Intelligence (cs.AI), new listings for Tuesday, 1 September 2026 (783 entries) — https://arxiv.org/list/cs.AI/new (accessed 2026-09-01)
- Hua, Yang, Hao, Zhang, Cao, Qi, Li, Guo — Agentic Quantitative Trading: A Survey of Workflows, Systems, and Evaluation (submitted 31 Aug 2026; announced 1 Sep 2026) — https://arxiv.org/abs/2608.31041 (accessed 2026-09-01; listed on [2])
- Zhao, Chen, Su — PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management (v4, revised 31 Aug 2026) — https://arxiv.org/abs/2605.27887 (accessed 2026-09-01; listed on [3])
- Kumar — Deep Hedging Under Realistic Market Frictions: A Regime-Conditional Empirical Study of Dynamic Option Hedging on Bitcoin Options (submitted 29 Aug 2026) — https://arxiv.org/abs/2608.29025 (accessed 2026-09-01; listed on [4])
- Bongiorno, Villassero — End-to-End Neural Shrinkage of Indefinite Pairwise Correlation Matrices for Small-Cap-Inclusive Portfolios — https://arxiv.org/abs/2608.30446 (accessed 2026-09-01; listed on [3])
- Angulo, Yeste, Espinos-Morato — Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation (submitted 29 Aug 2026) — https://arxiv.org/abs/2608.29463 (accessed 2026-09-01; listed on [5])
- Le, Tan, Williams-King — Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities — https://arxiv.org/abs/2608.29458 (accessed 2026-09-01; listed on [5])
- Chen, Zhao, Fu, Liang, Wu, Li, Xue, Zeng, Zhen, Xu, Li — Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment — https://arxiv.org/abs/2608.29696 (accessed 2026-09-01; listed on [5])
- Zhao, Shi, Xiao, Liu, Li, Hao, Hou, Guo, Zhang, Zhao, Wang, Liu, Wu, Yang, Sun, Zhang — ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions (EMNLP 2026 Main) — https://arxiv.org/abs/2608.30517 (accessed 2026-09-01; listed on [5])
- Goliath, Gebbie — Metaorder modelling and identification from public data (v1 23 Feb 2026; v2 revised 3 Sep 2026) — https://arxiv.org/abs/2602.19590 (accessed 2026-09-04; re-listed on [1] as a replacement for Friday, 4 September 2026. This note originally cited arXiv:2608.30999, which the authors' v2 comment records as the same version "submitted as a new work by accident"; the identifier was corrected on 2026-09-04 and the finding is unchanged.)
- Bundi — Optimal Block Time for AMM Liquidity Providers under Jump-Diffusion Prices (MARBLE 2026 extended version) — https://arxiv.org/abs/2608.30321 (accessed 2026-09-01; listed on [1])
- Nutz, Voss — The Convergence Rate of Stochastic Tracking with Application to Optimal Execution — https://arxiv.org/abs/2608.29468 (accessed 2026-09-01; listed on [1])
- Zhu, Pan, Liu, Hu, Wu — AI Can Be Easily Persuaded in Clinical Decision Making — https://arxiv.org/abs/2608.29453 (accessed 2026-09-01; listed on [5])
- Molent, Vellekoop — Neural Calibration of a Complete Market Model — https://arxiv.org/abs/2608.30867 (accessed 2026-09-01; listed on [2])
- Anthropic — Newsroom (dates the Fable 5.1 / Mythos 5.1 announcement to 1 September 2026) — https://www.anthropic.com/news (accessed 2026-09-04)
- Anthropic — Introducing Claude Fable 5.1 and Claude Mythos 5.1 — https://www.anthropic.com/claude-fable-and-mythos-5-1 (accessed 2026-09-04; dated on [19])
- OpenAI — Newsroom (dates the GPT-6 Astra system card and safety overview to 3 September 2026) and Path to Astra: critical capabilities and frontier safeguards, 1 September 2026 — https://openai.com/news/ and https://openai.com/index/path-to-astra/ (both accessed 2026-09-04)
- OpenAI — GPT-6 Astra: A new generation of intelligence — https://openai.com/index/gpt-6-astra/ (accessed 2026-09-04; dated on [21])
- arXiv — Portfolio Management (q-fin.PM) and Trading and Market Microstructure (q-fin.TR), new listings for Friday, 4 September 2026 (4 entries each) — https://arxiv.org/list/q-fin.PM/new (accessed 2026-09-04)
- Asaad, Mohamed, Zhang, Abdelsalam — The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis (submitted 2 Sep 2026) — https://arxiv.org/abs/2609.03218 (accessed 2026-09-04; listed on [23])