Skip to content

Commit ab91339

Browse files
anidotnetclaude
andcommitted
Take a sorted page's order from the index instead of every document
`find(ALL, orderBy("createdAt", Descending).limit(20))` asked for 20 rows and cost what draining the whole collection costs. `SortedDocumentStream` collects the entire result set before `BoundedStream` gets to drop 99% of it, and the cost is the decode, not the comparison, so it scales with document *size* as much as count. An index on the sort field bought nothing — the index was only ever used to filter — and page 50 cost exactly what page 1 cost, because the work finished before the skip applied. An index on the sort field already stores that field's value for every document. When the query has no filter, one sort field, a limit, and a simple unique or non-unique index on exactly that field, the keys now come from the index (`NitriteIndex.readSortKeys`) and `IndexedStream` fetches only the rows the page returns. On MVStore over 2000 rows each carrying a 150-element list, a `limit(20)` page went from ~125 ms to ~1 ms. The index replaces the documents only when it is a faithful stand-in for them: a multi-valued field is indexed once per element, so a duplicate-id check and an entry-count check catch it and fall back to the blocking sort. Ordering is decided by the same comparator either way — extracted as `DocumentSorter.compareValues` so the two cannot drift — over keys read from the index rather than from documents, which keeps nulls first and resolves ties in document-id order exactly as before. Where documents are small the index walk replaces a decode that was nearly free, so a sorted page over lean rows can cost a few hundred microseconds more than it did. The trade is deliberate: that loss is bounded and sub-millisecond, the win grows without bound with document size. New API is additive. `NitriteIndex.readSortKeys(long)` and `NitriteIndexer.readSortKeys(IndexDescriptor, NitriteConfig, long)` are `default` methods returning `null`, so an existing indexer plugin is unaffected and simply never takes the new path. `CollectionSortedFindTest` pins every sorted query against the same query on an unindexed collection — ascending, descending, deep pages, ties, strings, missing fields, unique indexes, and the multi-valued fallback — so a divergence in ordering fails outright. `CollectionSortedFindCostTest` runs on MVStore, where documents are actually serialized, and compares a sorted page over lean documents against one over fat documents at the same row count: only the decode differs between them, so it is stable under load in a way a wall-clock threshold is not (it reports 94x without this change and ~1x with it). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 parent fa94cb7 commit ab91339

23 files changed

Lines changed: 596 additions & 51 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,14 @@
1+
## Release 5.1.0 - Aug 14, 2026
2+
3+
### Improvements
4+
5+
- A sorted, limited `find` no longer fetches the whole collection when the sort field is indexed
6+
- `find(ALL, orderBy("createdAt", Descending).limit(20))` asked for 20 rows and cost what draining every stored document costs. `SortedDocumentStream` collects the entire result set before `BoundedStream` gets to drop 99% of it — and the cost is the decode, not the comparison, so it scales with document *size* as well as count. An index on the sort field bought nothing: the index was only ever used to *filter*, never to order, and page 50 cost exactly what page 1 cost because the work finished before the skip applied.
7+
- When the query has no filter, one sort field, a limit, and a simple unique or non-unique index on exactly that field, the sort keys are now read from that index — which already stores them — and only the documents actually returned are fetched. Measured over 2000 rows each carrying a 150-element list, a `limit(20)` page went from ~125 ms to ~1 ms, and stopped growing with the collection.
8+
- The index is used only when it holds exactly one entry per stored document. A multi-valued field is indexed once per element, which is detected by a duplicate-id check and an entry-count check, and falls back to the blocking sort. Ordering — including where nulls sort and how ties break — is identical either way: the same comparator runs over keys taken from the index instead of from the documents.
9+
- Where documents are small and cheap to read, the index walk replaces a decode that was nearly free, so a sorted page over lean rows can cost a few hundred microseconds more than before. The trade is deliberate: the loss is bounded and sub-millisecond, the win grows without bound with document size.
10+
- New API, all additive: `FindPlan.getSortIndexDescriptor()`, `NitriteIndex.readSortKeys(long)` and `NitriteIndexer.readSortKeys(IndexDescriptor, NitriteConfig, long)` (both `default`-implemented to return `null`, so existing indexer plugins are unaffected), and `DocumentSorter.compareValues(Object, Object, Collator)`.
11+
112
## Release 5.0.0 - Aug 7, 2026
213

314
### Breaking Changes

‎nitrite-bom/pom.xml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
<parent>
55
<groupId>org.dizitart</groupId>
66
<artifactId>nitrite-java</artifactId>
7-
<version>5.0.0</version>
7+
<version>5.1.0</version>
88
</parent>
99

1010
<artifactId>nitrite-bom</artifactId>

‎nitrite-jackson-mapper/pom.xml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
<parent>
55
<groupId>org.dizitart</groupId>
66
<artifactId>nitrite-java</artifactId>
7-
<version>5.0.0</version>
7+
<version>5.1.0</version>
88
</parent>
99

1010
<artifactId>nitrite-jackson-mapper</artifactId>

‎nitrite-mvstore-adapter/pom.xml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
<parent>
55
<groupId>org.dizitart</groupId>
66
<artifactId>nitrite-java</artifactId>
7-
<version>5.0.0</version>
7+
<version>5.1.0</version>
88
</parent>
99

1010
<artifactId>nitrite-mvstore-adapter</artifactId>
Lines changed: 125 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,125 @@
1+
/*
2+
* Copyright (c) 2017-2021 Nitrite author or authors.
3+
*
4+
* Licensed under the Apache License, Version 2.0 (the "License");
5+
* you may not use this file except in compliance with the License.
6+
* You may obtain a copy of the License at
7+
*
8+
* http://www.apache.org/licenses/LICENSE-2.0
9+
*
10+
* Unless required by applicable law or agreed to in writing, software
11+
* distributed under the License is distributed on an "AS IS" BASIS,
12+
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
13+
* See the License for the specific language governing permissions and
14+
* limitations under the License.
15+
*
16+
*/
17+
18+
package org.dizitart.no2.integration.collection;
19+
20+
import org.dizitart.no2.Nitrite;
21+
import org.dizitart.no2.collection.Document;
22+
import org.dizitart.no2.collection.FindOptions;
23+
import org.dizitart.no2.collection.NitriteCollection;
24+
import org.dizitart.no2.common.SortOrder;
25+
import org.dizitart.no2.index.IndexOptions;
26+
import org.dizitart.no2.index.IndexType;
27+
import org.dizitart.no2.mvstore.MVStoreModule;
28+
import org.junit.After;
29+
import org.junit.Before;
30+
import org.junit.Test;
31+
32+
import java.util.ArrayList;
33+
import java.util.List;
34+
35+
import static org.dizitart.no2.filters.Filter.ALL;
36+
import static org.dizitart.no2.integration.TestUtil.deleteDb;
37+
import static org.dizitart.no2.integration.TestUtil.getRandomTempDbFile;
38+
import static org.junit.Assert.assertTrue;
39+
40+
/**
41+
* The cost half of {@link CollectionSortedFindTest}, which needs a store that actually
42+
* serializes documents.
43+
* <p>
44+
* A blocking sort deserializes every stored document to read one field, so
45+
* {@code orderBy(indexed).limit(20)} cost what draining the whole collection cost - and the
46+
* gap grows with document size, not just document count. Taking the sort keys from the index
47+
* removes the decode: over 2000 rows carrying a 150-element array the sorted page went from
48+
* ~37 ms to ~1 ms.
49+
*
50+
* @author Anindya Chatterjee
51+
*/
52+
public class CollectionSortedFindCostTest {
53+
private static final int ROWS = 2000;
54+
55+
private final String fileName = getRandomTempDbFile();
56+
private Nitrite db;
57+
58+
@Before
59+
public void setUp() {
60+
db = Nitrite.builder()
61+
.loadModule(MVStoreModule.withConfig().filePath(fileName).build())
62+
.openOrCreate();
63+
}
64+
65+
@After
66+
public void tearDown() {
67+
if (db != null && !db.isClosed()) {
68+
db.close();
69+
}
70+
deleteDb(fileName);
71+
}
72+
73+
/**
74+
* The same query, the same row count, the same index - only the size of the documents
75+
* differs. A sorted page that decodes every row pays for the payload of every row, so
76+
* the fat collection costs many times the lean one. A sorted page that decodes only the
77+
* rows it returns costs the same either way.
78+
* <p>
79+
* Deliberately not "sorted page vs. full drain": both halves here are measured back to
80+
* back under whatever load the machine is under, and the only variable between them is
81+
* the thing that was broken.
82+
*/
83+
@Test
84+
public void testSortedPageCostDoesNotFollowDocumentSize() {
85+
double lean = sortedPageCost("lean", 0);
86+
double fat = sortedPageCost("fat", 150);
87+
88+
assertTrue("a sorted page over fat documents took " + fat + "ms against " + lean
89+
+ "ms over lean ones, same row count - it is still decoding rows it discards",
90+
fat < lean * 4);
91+
}
92+
93+
private double sortedPageCost(String name, int payloadSize) {
94+
NitriteCollection collection = db.getCollection(name);
95+
collection.createIndex(IndexOptions.indexOptions(IndexType.NON_UNIQUE), "seq");
96+
97+
for (int i = 0; i < ROWS; i++) {
98+
Document document = Document.createDocument("seq", i);
99+
if (payloadSize > 0) {
100+
List<Document> payload = new ArrayList<>(payloadSize);
101+
for (int w = 0; w < payloadSize; w++) {
102+
payload.add(Document.createDocument("text", "word" + w).put("start", w * 300));
103+
}
104+
document.put("payload", payload);
105+
}
106+
collection.insert(document);
107+
}
108+
109+
FindOptions page = FindOptions.orderBy("seq", SortOrder.Descending).limit(20);
110+
return timeOf(() -> {
111+
for (Document ignored : collection.find(ALL, page)) {
112+
// force the fetch
113+
}
114+
});
115+
}
116+
117+
private static double timeOf(Runnable task) {
118+
task.run(); // warm
119+
long start = System.nanoTime();
120+
for (int i = 0; i < 3; i++) {
121+
task.run();
122+
}
123+
return (System.nanoTime() - start) / 3e6;
124+
}
125+
}

‎nitrite-native-tests/pom.xml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@
77
<parent>
88
<groupId>org.dizitart</groupId>
99
<artifactId>nitrite-java</artifactId>
10-
<version>5.0.0</version>
10+
<version>5.1.0</version>
1111
</parent>
1212

1313
<artifactId>nitrite-native-tests</artifactId>

‎nitrite-rocksdb-adapter/pom.xml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
<parent>
55
<groupId>org.dizitart</groupId>
66
<artifactId>nitrite-java</artifactId>
7-
<version>5.0.0</version>
7+
<version>5.1.0</version>
88
</parent>
99

1010
<artifactId>nitrite-rocksdb-adapter</artifactId>

‎nitrite-spatial/pom.xml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
<parent>
55
<groupId>org.dizitart</groupId>
66
<artifactId>nitrite-java</artifactId>
7-
<version>5.0.0</version>
7+
<version>5.1.0</version>
88
</parent>
99

1010
<artifactId>nitrite-spatial</artifactId>

‎nitrite-support/pom.xml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
<parent>
55
<groupId>org.dizitart</groupId>
66
<artifactId>nitrite-java</artifactId>
7-
<version>5.0.0</version>
7+
<version>5.1.0</version>
88
</parent>
99

1010
<artifactId>nitrite-support</artifactId>

‎nitrite/pom.xml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
<parent>
55
<groupId>org.dizitart</groupId>
66
<artifactId>nitrite-java</artifactId>
7-
<version>5.0.0</version>
7+
<version>5.1.0</version>
88
</parent>
99

1010
<artifactId>nitrite</artifactId>

0 commit comments

Comments
 (0)