Michael McCandless

1.2K posts

Michael McCandless

Michael McCandless

@mikemccand

Apache member, PMC member and committer for the Apache Lucene project; Senior Principal Engineer at Amazon Product Search. Opinions are all my own!

Lexington, MA Katılım Ekim 2018
386 Takip Edilen1.1K Takipçiler
Michael McCandless
Michael McCandless@mikemccand·
OK in case anyone else has also wondered/struggled over the years why on earth they are called "daemons" in the UNIX world... I asked Claude. I hope it is not hallucinating; I like this origin story: claude.ai/share/923351b7…
English
0
0
1
182
Michael McCandless
Michael McCandless@mikemccand·
The 2025 Community Over Code NA Apache conference in Minneapolis was awesome! I recorded most of the talks from the Search track, here: youtube.com/playlist?list=… These are 360 videos! So you can slide/rotate your device to pick which direction to look.
English
1
0
1
322
Michael McCandless retweetledi
Adrien Grand
Adrien Grand@jpountz·
Lucene is getting an increasing number of high-quality contributions from ByteDance employees, especially around performance. Good to see that this project keeps attracting contributors from all around the world.
English
1
5
29
2.7K
Michael McCandless retweetledi
Adrien Grand
Adrien Grand@jpountz·
Another common point I did not expect: Vespa's strict vs. unstrict iterators is quite similar to Lucene's two-phase iteration. And both projects use this feature to effectively combine dynamic pruning with filtering (a hard and underappreciated problem IMO).
English
1
1
7
843
Michael McCandless retweetledi
Adrien Grand
Adrien Grand@jpountz·
There has been a big regression in Lucene's nightly benchmarks recently after a kernel upgrade. @mikemccand and @rcmuir found that it was caused by a change in the Linux scheduler configuration. #issuecomment-2889212582" target="_blank" rel="nofollow noopener">github.com/apache/lucene/…
English
1
3
6
672
Michael McCandless retweetledi
Adrien Grand
Adrien Grand@jpountz·
Yelp's nrtSearch was just upgraded to Lucene 10. Also switched from persistent storage to object storage as a source of truth, and plans on doing NRT replication via object storage instead of over the network. Very similar to Elasticsearch Serverless.
Yelp Engineering@YelpEngineering

In our new post, Sarthak and Andrew walk us through the big updates in Nrtsearch V1! Discover how incremental backups, smarter state management, ephemeral local disks, and a move to Lucene 10 are powering faster, more reliable search at Yelp. 🔗: bit.ly/3YACAp4

English
0
4
13
921
Michael McCandless retweetledi
Adrien Grand
Adrien Grand@jpountz·
The search library benchmark from the Tantivy folks was just updated with Lucene 10.2 tantivy-search.github.io/bench. Lucene now performs much better at the COUNT collection type, a bit better at TOP_K. Still somewhat slow at TOP_100_COUNT and phrase queries across all collection types.
English
0
2
16
675
Michael McCandless retweetledi
Cory Zue
Cory Zue@czue·
This xkcd was published 18 years ago, mostly irrelevant for like a decade as computers got faster, and is now very much back.
Cory Zue tweet media
English
2
1
7
670
Michael McCandless retweetledi
Andy Jassy
Andy Jassy@ajassy·
Important moment for @ProjectKuiper as we just confirmed our first 27 production satellites are operating as expected in low Earth orbit. While this is the first step in a much longer journey to launch the rest of our low Earth orbit constellation, it represents an incredible amount of invention and hard work. Am really proud of the collective team.
Andy Jassy tweet media
English
51
104
669
241.1K
Michael McCandless retweetledi
Adrien Grand
Adrien Grand@jpountz·
Guo Feng contributed a 2.5x (!) speedup to #Lucene's numeric range queries by using vectorization. HZ sped up query evaluation, ID sped up decoding data from the index. Lots of great performance improvements coming in Lucene 10.2.
Adrien Grand tweet media
English
0
6
26
1K
Michael McCandless retweetledi
Adrien Grand
Adrien Grand@jpountz·
This is now live on nightly benchmarks, with a 32% speedup on primary-key lookups benchmarks.mikemccandless.com/PKLookup.html and a 5% speedup on fuzzy queries. The main benefit that Lucene users will notice is likely faster indexing with explicity document IDs
Michael McCandless@mikemccand

#Apache #Lucene will soon have a faster and smaller terms index! This is a complex part of Lucene, and a major hotspot for terms heavy use cases like (primary) key/value store (~34% speedup, but results are preliminary!). Lucene's pluggable Codec API makes experimentation like this possible. #issue-2899660205" target="_blank" rel="nofollow noopener">github.com/apache/lucene/…

English
0
2
11
652
Michael McCandless retweetledi
Adrien Grand
Adrien Grand@jpountz·
Lucene's histogram collector is becoming more sophisticated, it can now take advantage of points indexes when the query fully matches a segment, which can give a multiple fold performance boost. github.com/apache/lucene/…
English
0
2
9
533
Michael McCandless
Michael McCandless@mikemccand·
#Apache #Lucene will soon have a faster and smaller terms index! This is a complex part of Lucene, and a major hotspot for terms heavy use cases like (primary) key/value store (~34% speedup, but results are preliminary!). Lucene's pluggable Codec API makes experimentation like this possible. #issue-2899660205" target="_blank" rel="nofollow noopener">github.com/apache/lucene/…
English
1
5
28
1.5K
Michael McCandless
Michael McCandless@mikemccand·
I must call out this unsung hero in #Apache #Lucene's benchmarking toolkit: Created by @rcmuir (thank you!), it tests performance impact of any git branch (or PR) against many (10 currently?) different CPU architectures/models/revisions via AWS's diverse instances so we can catch accidental regressions in our SIMD performance. You just run a single "make" command. It's because of passionate developers and modern cloud-native tools like these that #Lucene keeps improving so quickly! And that's so hard in this modern brittle world of tapping into vectorized (SIMD) instructions in modern CPUs from all the way up the clouds of pure avaland. github.com/apache/lucene/…
English
2
2
22
960
Michael McCandless
Michael McCandless@mikemccand·
The next Apache #Lucene major release (11.0) already has some compelling performance gains over 10.x, now encoding dense postings blocks as a bitset, and using vectorized (SIMD) CPU instructions for postings int[] decoding! Thanks to the pluggable Codec API in Lucene, such experimentation and improvement is (relatively) simple! #issuecomment-2589415510" target="_blank" rel="nofollow noopener">github.com/apache/lucene/…
English
2
3
18
909