Jane Street研究实习生Kavish开发了一个基于自回归扩散的市场数据生成模型1。该模型针对市场订单簿事件及其发生时间进行合成,旨在探索如何处理市场数据兼具连续性和离散性的特点1。
研究过程中采用了四年美国股票数据进行训练1。在模型对比中发现,DDPM方法在高噪声水平下性能不稳定,其中88-95%的生成值偏离分布均值8个标准差,而流匹配方法相比之下表现显著更好1。为了处理数据分布中的不连续性,研究团队采用了"原子平滑"技术,并通过20分类的离散头来区分不同的市场事件类型1。
评估方面,研究采用总变差散度来衡量生成分布与真实分布的偏差,同时利用分类器来区分真实事件和生成事件以评估样本质量1。最终模型能够生成相对逼真的市场数据样本,但精度尚不足以用于实际应用1。
Kavish, a research intern at Jane Street, has developed a generative model based on autoregressive diffusion to synthesize order book events and their timing in financial markets1. The project addresses the dual nature of market data, which contains both continuous and discrete components, by implementing flow matching techniques that significantly outperformed DDPM approaches1. When tested, DDPM exhibited instability at high noise levels, with 88-95% of generated values deviating more than eight standard deviations from the distribution mean1.
To handle the discrete aspects of market events, the final model employs a 20-category discrete head to classify different types of order book activities1. The researchers applied "atomic smoothing" to mitigate discontinuities inherent in the data distribution1. Model evaluation relied on total variation divergence to measure discrepancies between generated and real distributions, as well as classifier-based assessment to distinguish authentic events from synthetic ones1. The training utilized four years of U.S. stock market data1.
While the model demonstrates the feasibility of generating relatively realistic market data samples, the accuracy remains insufficient for practical trading applications1.
评论
还没有评论,欢迎留下第一条。