JetBrains公开了其语义代码搜索平台Air Context的设计与实现方案1。该系统通过结构感知的代码分块、向量化处理和二进制量化等技术,为AI代理提供精确的代码检索能力1。
Air Context的核心技术包括从代码解析、智能分块到向量存储的完整RAG管道1。该平台支持Kotlin、Java、Python、JavaScript、TypeScript、C#、PHP、Go和Rust等语言的解析和结构感知分块1。为了大幅降低存储成本,系统采用二进制量化技术将向量从32位浮点数压缩为1位,使存储空间缩小32倍,同时使用汉明距离替代余弦相似度进行向量比较1。
隐私保护与性能的平衡是该方案的重要特性1。所有代码索引使用开源权重嵌入模型在JetBrains自有GPU上运行,不依赖第三方云模型,以保护客户隐私1。查询结果以坐标形式返回,代码片段在用户本地从源代码检出组装,服务器不存储源代码内容1。在实际应用中,JetBrains的IntelliJ IDEA单体仓库包含超过100万个文件,最长路径达218字符1。二进制向量相关得分范围被压缩至较窄区间,影响阈值过滤的准确性,某些场景需保留16位浮点精度1。
JetBrains has shared the technical architecture and implementation experience of Air Context, its semantic code search platform, detailing how the system leverages structure-aware code chunking, vectorization, and binary quantization to deliver precise code retrieval capabilities for AI agents 1. The platform supports parsing and structure-aware chunking across nine programming languages: Kotlin, Java, Python, JavaScript, TypeScript, C#, PHP, Go, and Rust 1.
The RAG pipeline employs binary quantization technology to compress vector sizes from 32-bit floating-point numbers to 1-bit representations, reducing storage space by 32 times while using Hamming distance instead of cosine similarity for comparisons 1. All code indexing utilizes open-source embedding models running on JetBrains' own GPUs rather than third-party cloud models, ensuring customer privacy protection 1. Query results are returned as coordinates, with code snippets assembled locally from source code on users' machines rather than stored on servers 1.
The IntelliJ IDEA monolithic repository contains over one million files with paths reaching up to 218 characters 1. Binary vector relevance scores are compressed into a narrow range, which can affect the accuracy of threshold filtering in certain scenarios, sometimes necessitating the retention of 16-bit floating-point precision 1.
评论
还没有评论,欢迎留下第一条。