Skip to content
Preprint

Point-in-Time Audit Before Alpha: Public-Archive Availability and a Negative Matched-Budget Study on BTC Perpetual Futures

Aug 2026 · 0 citations · 15 references
Computer Science

Abstract

Public cryptocurrency archives may appear usable when files exist, although factor research requires observations available and executable at each decision time. We audit public Binance BTCUSDT USD-M perpetual-futures data using event, publication, and availability times and separate proposal from deterministic auditing, evaluation, and holdout access. An initial gapless five-minute requirement for trade, mark, index, and open interest failed: the longest unrepaired intersection was 304.5729166666667 days. A disclosed revision made trade, mark, index, and realized funding the core streams and made open interest optional because its publication time was unverified. The revised mask retained 727 complete UTC days and supported a 436/145/146-day train, validation, and historical-holdout split. On 80 frozen known-rule templates, the auditor detected 40/40 violations and rejected 0/40 legal templates. Across ten null-signal paths, full auditing reduced mean false passes from 0.2910 to 0.0625. Under matched valid-candidate budgets, the audited adaptive agent tied random search and did not establish superiority. In the one-time historical holdout, all evaluated runs had positive IC but negative net Sharpe under primary costs. We therefore report a scoped negative result rather than a profitability or agent-superiority claim.

View source

Similar papers

Preprint Aug 2026

OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents

This work presents OpenPM, an auditable point-in-time evaluation framework for LLM portfolio-management agents, and builds a reference agent named the tiered allocator, where typed analysts score candidates, a constructor LLM proposes weights, and a deterministic critic guarantees feasibility.

Xinying Cai, Ming-Hao Guo, Jiahe Liu et al. · 0 citations
Preprint Jul 2026

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked purposes. First, it audits a maximum-variation purposive corpus of 40 empirical records appearing between 18 July 2025 and 17 July 2026. The audit coded publication route, execution timing, model identity, age of the newest named generation or immutable snapshot, same-family supersession and refresh behaviour. At appearance, the newest named model was a median 281 days old (middle 50%: 75-478; range: 11-939). Median age was 395 days for 25 journal articles, 56 days for 14 preprints and 49 days for one laboratory report. Thirty-five records included a superseded family, seven supplied a precise dated identifier, three clearly refreshed model evidence, and one added a late sensitivity test. All 40 included an OpenAI system, a feature of this corpus rather than a prevalence estimate. The paper distinguishes model age from claim currency and proposes six reporting practices. Second, it treats its own two-day production process as a reflexive case of frontier-model-assisted research creation. GPT-5.6 Sol Pro in ChatGPT supported candidate discovery, source reconciliation, calculations, drafting and critique; the author checked sources, made all substantive decisions and accepts responsibility. This is a proof-of-practice, not a controlled estimate of productivity or quality. By applying its own Model Facts and model-currency statement, the paper shows how rapid AI-assisted research can be made inspectable without treating model output as independent validation. The title uses half-lives metaphorically; no universal decay rate is estimated.

C. Iacono · 0 citations
Preprint Aug 2026

The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence

The resulting research prototype binds each deterministic policy decision to the exact policy source, commits a privacy-minimizing record at a caller-selected synchronization boundary, and returns an Ed25519-signed receipt that states whether that boundary completed.

Neeraj Kumar Singh Beshane · 0 citations
Open access Aug 2026

Sealed Before the Event: A Publicly Verifiable, Bitcoin-Anchored Forecast Record of 92 Graded Outcomes, with a Source-Agnostic Protocol for Scoring Unexplainable Forecast Sources

Background Public prediction records are hard to evaluate: posts can be edited or deleted, and even honest records rarely separate anteriority (did this exact claim exist unchanged before the event?) from skill (do the claims carry predictive information?). A motivating case: a public forecast naming mass-casualty attacks on public gatherings in Moscow in a March–April 2024 window was published and server-timestamped on 4 September 2023 — ~200 days before the Crocus City Hall attack and 185 days before a comparable official alert. Methods Every forecast was published verbatim before its window; each text was canonicalised, SHA-256-hashed into a manifest, and anchored to the Bitcoin blockchain via OpenTimestamps — the anchor fixes integrity (the bytes unchanged since the anchoring block), while anteriority rests on platform-issued timestamps, pre-event public distribution, and third-party archives (§2.1.2); a four-level rubric (HIT/NEAR/PARTIAL/MISS) was fixed at seal time and misses retained. We formalise a five-step source-agnostic scoring protocol for sources whose generative mechanism is unavailable or unexplained: demand cryptographic anteriority; freeze the claim verbatim; grade the stream, never the anecdote; set weight by calibration, not theory; and let the frozen denominator filter uninformative sources. Results Across 92 graded forecasts (73 HIT, 6 NEAR, 4 PARTIAL, 9 MISS) the self-assigned aggregate Brier is 0.0958 (launch subset 0.036; strategic warning 0.116). On this small, high-confidence sample a base-rate baseline ties the aggregate Brier — conceded up front. A pre-specified clustered luck test — correlated calls collapsed into independent events, strict all-HIT scoring, luck prior floored at 0.5 per event — yields 51 of 68 strict successes (exact binomial p = 2.2 × 10 −5 ; break-even luck prior ≈ 0.65). A five-way calibration-integrity suite is reported with sample-size ceilings. All statistics recompute from public artifacts with zero-dependency tooling. Conclusions The contribution is the protocol — belief-independent evaluation of unexplainable forecast sources — not validation of the disclosed generative method (Vedic jyotish, treated as a black box). Applications to warning analysis and war-game parameterization are outlined; the limitations — self-grading, sample size, operator selection — are load-bearing.

Vijay Jyotish · 0 citations
Review Aug 2026

KONTOGRAPH: Verified Point-in-Time Feature Consistency and Amortised Explanation for Real-Time Anti-Money Laundering under a 200 ms Decision Budget

Regulation (EU) 2024/886 obliges European payment service providers to settle euro credit transfers in under ten seconds, around the clock. This removes both the overnight batch window in which anti-money-laundering (AML) analytics traditionally ran and the settlement delay that made recovery possible, forcing detection, explanation and decision inside a single-digit-second envelope. We present KONTOGRAPH, an end-to-end AML pipeline for the SEPA Instant rail built under a self-imposed 200 ms 99th-percentile budget, and report an empirical study on 1,562,860 simulated payments with injected typologies and deliberately incomplete labels. Three findings are of interest beyond the system itself. First, a temporal graph network with per-node memory improves PR-AUC over a gradient-boosted tabular baseline from 0.0053 to 0.1717, a paired day-blocked bootstrap difference of +0.166 with 95% CI [0.105, 0.241]; per-node memory alone more than doubles the score. Second, expressing each feature once and compiling it to three execution backends, with equivalence enforced by property-based tests that perturb the future, surfaced three point-in-time violations that code review had passed--each of which would have inflated reported performance. Third, and most consequential for practice, exporting the deployed tree ensemble to ONNX changed only $7.4 \times 10^{-8}$ in mean score yet altered 0.26% of decisions and inflated the alert volume by 12%, because 32-bit accumulation perturbs scores across a cost-optimal threshold of $3.98 \times 10^{-4}$. We argue that a serving-format conversion must be treated as a model change until measured, and that fidelity metrics for subgraph explainers can be vacuous when candidate neighbourhoods are small--a null result we report in full.

Ahmed Abolfadl · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.