Paper: arXiv 2609.18975

Authors: Arati Uday Kamat

Abstract

On-chain studies of memecoin launchpads usually rely on one data-collection pipeline, yet whether differently configured pipelines observe the same tokens is rarely measured. This paper compares the output mint sets of two separately configured pipelines from one research programme on the Solana pump.fun launchpad: a cohort-detection pipeline that flags coordinated early-buyer wallets, and a rejection-filtering pipeline that logs a trader-side observer’s pre-trade filter decisions. Two consecutive, non-overlapping windows are analysed (v1: June 2026; v2: June-July 2026); in each, the rejection data are restricted to the interval spanned by the cohort detections under four timestamp rules. Only raw counts, observed proportions and Jaccard/Dice indices are reported. In v2, where the rejection stream covers the whole 17.5-day window, the cohort set contains 623 mints and the rejection set 1,742, with 5 overlapping mints (0.803% of cohort; 0.287% of rejection). In v1 a collection gap means the rejection stream covers only the final 14.25 hours (4.4%) of the 13.4-day cohort window; the 20,162 cohort and 53 rejection mints share 1 mint, and within the covered interval none, a count too small to be informative. Both windows show nearly disjoint outputs, agreeing in direction but not in magnitude. Collector configuration, not only market behaviour, therefore shapes which tokens an on-chain study observes, and the paper proposes reporting cross-collector overlap and collector coverage as a routine check. Data and a script that re-derives every number are openly available.

Complexity vs Empirical Score

  • Math Complexity: 2.0/10
  • Empirical Rigor: 8.0/10
  • Quadrant: Street Traders — practical and empirical, lighter on theory

Why this score: This paper presents a highly rigorous empirical comparison of on-chain data pipelines with strong reproducibility. While the mathematical complexity is low, the novel finding of pipeline-dependent token observation is significant for on-chain research. The paper is well-written and thorough in its analysis and discussion.

Research Flowchart

  flowchart TD
    A[Research Question: Do different on-chain observation pipelines see the same tokens?] --> B{Methodology: Compare mint sets from two pipelines on Solana pump.fun};
    B --> C1[Input: Cohort-detection pipeline output (mint sets)];
    B --> C2[Input: Rejection-filtering pipeline output (mint sets)];
    C1 & C2 --> D[Process: Analyze two time windows (v1, v2) for overlapping mints];
    D --> E[Process: Calculate Jaccard/Dice indices, raw counts, observed proportions];
    E --> F{Key Finding: Pipelines show nearly disjoint outputs};
    F --> G[Outcome: Collector configuration significantly shapes observed tokens, not just market behavior];
    G --> H[Recommendation: Report cross-collector overlap and coverage as routine check];