cd /news/ai-infrastructure/i-tested-qdrant-quantization-and-mma… · home › topics › ai-infrastructure › article
[ARTICLE · art-141981] src=pub.towardsai.net ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

I Tested Qdrant Quantization and mmap on EV Charger Search

A hands-on test of Qdrant's scalar, binary, TurboQuant, and product quantization on 15,205 real Great Britain EV charging locations and a separate 100,000-point synthetic workload found that smaller vector copies cut memory but not all preserved the correct nearest neighbours, with a measured Float32-cold Qdrant container using 592.90 MiB against a theoretical 293 MiB for 100,000 768-dimension Float32 vectors. The test also compared keeping full-size vectors warmed in the operating system cache versus left on disk, and used payload filters plus rescoring against the original Float32 vectors to check final candidates. The author framed the work around whether a growing EV charger semantic-search service can stay affordable without degrading results.

by read21 min views1 publishedSep 29, 2026

When I started this project, I had a question in mind: if a charging network keeps adding stations and search records, how much memory will its vector search need? A small demo can fit on a laptop. A large service has to pay for every byte it keeps ready in RAM.

I used Qdrant to test four ways to make vectors smaller: scalar, binary, TurboQuant, and product quantization. I also tested where Qdrant keeps the full-size vectors: warmed in the operating system’s cache or left on disk until needed. The memory savings were real. But a smaller vector was not always a good search vector. Some methods kept the right neighbours; others lost too many of them.

This article follows one question from beginning to end: how can an EV search service keep a growing semantic-search collection affordable without quietly damaging its results? I start by showing how much space one vector needs. Then I follow a search through Qdrant: how it makes a smaller copy for fast searching, keeps large files on disk, removes incompatible chargers, and checks the final candidates against the original vectors.. The evidence comes from 15,205 real charging locations in Great Britain and a separate 100,000-point workload made from seeded, fictional charging scenarios. I keep those workloads clearly separated because a synthetic scale test is not the same thing as live EV telemetry.

Imagine a European charging network serving a driver asking for a compatible, high-power charger near a route. A search record might describe a station’s connectors and nearby road context. An embedding model turns that description into a vector, which is simply a long list of numbers. Qdrant compares those lists to find records with similar meaning.

This project uses 768 numbers per vector. Each original number is a Float32 value that takes 4 bytes. So one vector takes 768 × 4 = 3,072 bytes. One million vectors take 3,072,000,000 bytes, or about 2.86 GiB, before the search graph, filters, payloads, and the database process itself. Ten million would need about 28.6 GiB just for those raw vectors.

That calculation left me with two different problems. First, I needed a smaller representation for the fast candidate-search stage. Second, I needed control over where the original full-size vectors lived. I chose Qdrant because one collection can keep the original Float32 vectors, add Scalar, Binary, TurboQuant, or Product Quantization as a compact search copy, place individual structures in different memory tiers, apply hard payload filters, and rescore a shortlist against the originals. That gave me one system in which I could change the compression method or memory placement while keeping the data, queries, filters, and evaluation logic the same. Qdrant’s quantization guide is here: https://qdrant.tech/documentation/manage-data/quantization/.

The raw-vector arithmetic is simple enough to check in Python. It is useful because it stops a compression ratio from sounding like a promise about the whole server:

def vector_mib(points, dimensions, bits_per_number):    return points * dimensions * bits_per_number / 8 / 2**20for bits in (32, 8, 4, 2, 1):    print(bits, round(vector_mib(100_000, 768, bits), 2))

The calculation above describes only the numbers inside the vectors. At 100,000 points, 768 Float32 values per point occupy about 293 MiB. Replacing each 32-bit value with 8, 4, 2, or 1 bit gives the idealized 73, 37, 18, and 9 MiB figures. A running Qdrant container needs more than that: it also holds the HNSW graph, payload and filter indexes, allocator and process overhead, file-backed pages in the operating-system cache, and — in a quantized collection — the original Float32 vectors used for rescoring. That is why “32× compression” is a statement about one compact vector copy, while the measured Float32-cold container in this run used 592.90 MiB. The theoretical calculation is a useful lower-level check; the container measurement is the number that describes this experiment.

Later, Figure 1 compares these theoretical compression figures with the container memory measured during the 100,000-point benchmark.

I used two public data sources. The charging locations came from the OpenChargeMap export at https://github.com/openchargemap/ocm-export. It supplied station coordinates, operator names, and connector information. The road context came from the UK Department for Transport’s GB road-traffic counts at https://www.data.gov.uk/dataset/208c0e7b-353f-4e2d-8b7a-1a7118467acc/gb-road-traffic-counts. Its AADF field estimates how many vehicles pass a road count point on an average day. I joined each charger to a nearby count point to give the station description some road context. AADF is an annual traffic context; it is not a live feed, charger occupancy, or a count of charging sessions.

I matched each station to its nearest DfT count point if it was within 5 km. That left 15,205 real charging locations. The source and licence details stay with the prepared records. For the larger run, the code made 100,000 repeatable scenarios by adding fictional battery percentage, arrival time, tariff, waiting time, and trip purpose. Those added fields are labelled synthetic. They make a larger technical workload, not a claim about real EV behaviour.

FastEmbed is the lightweight Python library that runs the embedding model locally. The model, BAAI/bge-base-en-v1.5, reads each station description and returns 768 numbers. You can think of those numbers as coordinates on a semantic map: descriptions with similar meaning should land near one another, even when they do not use exactly the same words. That fuzzy similarity is useful for questions such as “find a rapid charger for a long trip.” It is not trustworthy enough for a hard safety or compatibility rule. I therefore kept latitude, longitude, connector type, and charging power as structured Qdrant payload fields. A query may use the vector to rank relevant stations, while a payload filter separately requires, for example, that one connector is Type 2 and at least 50 kW. The embedding suggests; the filter enforces.

The diagram helped me keep the data story honest. OpenChargeMap gave me station locations and connector details. DfT gave me annual road counts near those locations. The join adds context about a nearby road; it does not measure how many drivers used a charger. The larger branch repeats real stations with seeded, made-up trip conditions so I can stress the search system.

Each record has two jobs. Its readable description becomes the 768-number vector. Its structured fields stay available for exact rules such as “the same connector must be Type 2 and offer at least 22 kW.” Here is a shortened version of the preparation step in the repository:

for record in records:    record["text"] = station_text(record)vectors = np.lib.format.open_memmap(    settings.prepared_dir / "vectors.npy",    mode="w+", dtype=np.float32,    shape=(len(records), settings.dimensions),)for i, vector in enumerate(    model.embed(r["text"] for r in records)):    vectors[i] = vectorvectors.flush()

The local vectors.npy file lets preparation write embeddings without holding the whole array in Python RAM. That is a separate use of mmap from Qdrant’s storage tiers. When the records are uploaded, Qdrant decides where its own original vectors and compact copies will live. I keep that distinction in mind because “on disk” can refer to different parts of the pipeline.

Think of a Float32 vector as a detailed picture. Quantization makes a smaller sketch that Qdrant can compare quickly. The sketch can be good enough to find candidates, but it may miss details. The four methods make that sketch in different ways.

Scalar quantization turns each 32-bit number into an 8-bit value. The compressed vector is nominally 4× smaller. It is the cautious starting point when I want a smaller search copy without a large accuracy drop. The actual loss depends on the model and data, so I measured it rather than treating “near-zero loss” as a guarantee.

One-bit binary quantization keeps only the sign of each vector coordinate: positive becomes one value and negative becomes the other. It forgets magnitude, so +0.02 and +3.00 look identical in that coordinate. The reward is a nominally 32× smaller copy and very cheap comparisons. The cost is that the method works best when the vector has many dimensions and its values are well balanced around zero, so the signs still carry enough information. Qdrant’s guide warns that one-bit compression can lose substantial precision below roughly one thousand dimensions. My BGE vectors have 768 dimensions, and the measured result exposed that risk: Binary 1-bit reached only 0.2605 Recall@10 with 2× rescoring. Binary quantization was doing what it was designed to do; the one-bit sketch simply threw away too much information for this model and workload.

TurboQuant first rotates the vector so that information is spread more evenly, then maps the rotated values into a small number of levels. The underlying technique can use fast Hadamard-style rotations and Lloyd–Max quantization levels. In plain English, it shuffles the picture before drawing a compact sketch so one awkward coordinate is less likely to dominate. Qdrant offers several bit depths; I tested 4-bit (nominally 8×) and 2-bit (nominally 16×). This does not require me to train a special embedding model.

Product quantization cuts a vector into chunks and stores a code for a representative pattern in each chunk. I tested its x32 setting. It can be useful when the compressed copy must be very small, but the code may be a rougher guide to the nearest neighbours. Qdrant still retains the original vectors, so this setting is not a magic answer to total disk cost.

Qdrant 1.19 introduced an explicit memory setting for each major collection structure; the documentation is at https://qdrant.tech/documentation/ops-configuration/memory-tiers/. The setting applies separately to dense vectors, the HNSW graph, quantized vectors, payloads, and payload indexes. Pinned means Qdrant allocates the structure on its heap and keeps it in RAM. Cached means the structure lives in a memory-mapped file and Qdrant asks the operating system to warm those pages at startup. Cold uses the same kind of memory-mapped file but skips that pre-warm, so the first access may require a disk read. Cached and cold pages can both enter or leave the OS page cache later. In this benchmark I pinned the small search structures — the HNSW graph, quantized copy, and numeric filter indexes — and left the much larger original vectors and payloads cold.

A memory-mapped file, usually called mmap, is a way to let a program treat parts of a disk file like memory. The operating system fetches pages when needed. This is why “on disk” does not mean “never uses RAM.” It also explains why restarting the Qdrant process does not necessarily give a truly cold disk measurement: the OS may still remember those pages.

In my quantized collections, the HNSW search graph, the small quantized vectors, and numeric filter indexes are pinned. The original Float32 vectors and the station payloads are cold. HNSW is simply the map Qdrant follows to avoid comparing the query with every vector. A separate Float32 cached collection and a Float32 cold collection show what changes when the original vector placement changes.

The new names replace several older switches, so the mapping depends on which structure you are configuring. For dense vectors, on_disk: false corresponds to memory: cached, while on_disk: true corresponds to memory: cold. The HNSW graph uses the same cached/cold mapping, although the new pinned tier has no old HNSW equivalent. For quantized vectors, the old always_ram: true corresponds to pinned. For payloads, on_disk_payload: false corresponds to cached and true corresponds to cold. Existing legacy configurations still work; Qdrant does not silently rewrite them during an upgrade. I used the Qdrant 1.19 memory names because memory: pinned, cached, or cold makes the intended location of each structure visible in the collection definition.

First, the filter removes incompatible stations. In the benchmark, connector type and minimum charging power are hard conditions. Qdrant then walks the HNSW graph and compares the query with compact vectors to gather likely matches. If rescoring is enabled, it reads the original Float32 vector for a smaller candidate group, recalculates their scores, and returns the best ten.

Here is the concrete version of oversampling. Suppose the user wants ten results. The compact quantized search first builds a shortlist. With no rescoring, Qdrant can return that approximate shortlist directly. With 1× rescoring, it checks roughly the same ten candidates against the original Float32 vectors; their order may improve, but a good station that never entered those ten cannot be recovered. With 2× oversampling, Qdrant keeps roughly twenty candidates, reads their original vectors, recalculates the scores, and returns the best ten. At 4× it does the same with roughly forty candidates. A larger shortlist gives a missed neighbour another chance to enter the final ten, but it also requires more full-vector reads and scoring work. In the 100,000-point scalar run, 1× did not improve Recall@10, while 2× raised it from 0.7095 to 0.7690; 4× added no further recall. That is the trade-off the oversampling setting controls.

A compact version of the Qdrant setting used in the project looks like this:

vectors_config = models.VectorParams(    size=768, distance=models.Distance.COSINE,    memory=models.Memory.COLD,)quantization_config = models.ScalarQuantization(    scalar=models.ScalarQuantizationConfig(        type=models.ScalarType.INT8,        memory=models.Memory.PINNED,    ))payload = models.PayloadStorageParams(memory=models.Memory.COLD)

The first code block creates a collection with original Float32 vectors cold and a pinned scalar copy. For the comparison, the rest of the collection stays the same: same 768 dimensions, cosine distance, HNSW settings, payload, filters, and points. I swap the quantization_config. This is a shortened view of two choices from the collection helper:

scalar = models.ScalarQuantization(    scalar=models.ScalarQuantizationConfig(        type=models.ScalarType.INT8,        quantile=0.99,        memory=models.Memory.PINNED,    ))turbo4 = models.TurboQuantization(    turbo=models.TurboQuantQuantizationConfig(        bits=models.TurboQuantBitSize.BITS4,        memory=models.Memory.PINNED,    ))The other two families use the same pinned-memory idea:binary = models.BinaryQuantization(    binary=models.BinaryQuantizationConfig(        encoding=models.BinaryQuantizationEncoding.ONE_BIT,        query_encoding=models.BinaryQuantizationQueryEncoding.BINARY,        memory=models.Memory.PINNED,    ))product = models.ProductQuantization(    product=models.ProductQuantizationConfig(        compression=models.CompressionRatio.X32,        memory=models.Memory.PINNED,    ))

For TurboQuant 2-bit, the helper changes BITS4 to BITS2. I tested each method in its own collection. “4×” or “32×” describes only the small vector representation; the full Float32 copy, HNSW graph, filter indexes, and payload still exist. That is why the measured container-memory reductions are much smaller than the ratios on the quantization ladder.

The actual repository builds every method through the same collection helper. The query setting changes search-time behaviour:

models.SearchParams(    hnsw_ef=128,    quantization=models.QuantizationSearchParams(        rescore=True, oversampling=2.0    ),)

That second line does not change the stored data. It tells Qdrant to gather extra candidates and check them against the original vectors for this query. The linked notebook prints all seven collection configurations and runs the same code path on a small artificial fixture.

One detail is easy to miss in the filter. A station can have several connectors. A loose filter might match the plug type on one connector and the power on another, then return a station with no single suitable plug. The benchmark uses Qdrant’s nested filter so both conditions must belong to the same connector:

plug_and_power = models.NestedCondition(    nested=models.Nested(        key="connectors",        filter=models.Filter(must=[            models.FieldCondition(                key="type_id",                match=models.MatchValue(value=33),            ),            models.FieldCondition(                key="power_kw",                range=models.Range(gte=50),            ),        ]),    ))

The query text still helps rank the eligible stations. The filter is the hard gate; a nice-sounding embedding cannot make an incompatible plug acceptable. This benchmark does not check live availability, route travel time, or whether a tariff is current. Those would need real service data and their own filters or ranking rules.

I froze 200 search questions and used the same questions, connector filters, and power filters for every collection. For each question, I also asked for the exact top ten with full Float32 vectors and quantization ignored. That exact list is the reference. Recall@10 asks: how many IDs from that exact list did the faster search return? If it finds eight of the ten, Recall@10 is 0.8. This is agreement with exact vector neighbours, not a human judgement that a charger recommendation is useful.

I measured three other things. P50 latency is the middle query time; P95 is a slower tail time that 95 percent of timed requests do not exceed. The memory report separates Qdrant’s estimated heap, cached pages, and the whole Docker container’s memory. Those numbers overlap, so I never add them together. Each method used an isolated Qdrant container, identical data and query settings, and saved per-query results so the totals can be recalculated.

The benchmark waited for indexing, persisted the collection, restarted it, verified every live point was still in indexed segments, warmed the queries, then timed repeated requests. That extra checking mattered: after one earlier restart, accepted points were not all ready for quantized search. Without the readiness check I could have reported a misleading result. The final published runs passed that check.

For each of the 200 frozen questions, I first ask Qdrant for the exact Float32 neighbours. The reference search deliberately ignores quantization. The important detail is using both settings together: exact=True alone could still leave the compressed search path involved in a quantized collection.

exact_params = models.SearchParams(    exact=True,    quantization=models.QuantizationSearchParams(        ignore=True    ),)reference_ids = [p.id for p in exact_points[:10]]candidate_ids = [p.id for p in fast_points[:10]]recall_at_10 = len(    set(reference_ids) & set(candidate_ids)) / 10

If eight IDs overlap, recall is 0.8. This says nothing about whether a human driver liked the suggestion. It tells me whether a cheaper search preserved the neighbours from the full-precision mathematical search.

I also sampled memory from two places. Qdrant’s collection endpoint estimates where the collection’s heap and cached pages sit. The container cgroup tells me the whole container’s memory at that moment. These are different views of overlapping bytes:

curl "$QDRANT_URL/collections/scalar_int8/memory"docker exec "$CONTAINER" cat /sys/fs/cgroup/memory.current

The report prints those views side by side. I did not add them together. I also separated warm queries from the first query after a restart, because a process restart does not promise an empty operating-system page cache.

On the 15,205 real station descriptions, the Float32 cold approximate baseline reached 0.9075 Recall@10 with 250.32 MiB of container memory. Scalar int8 with 2× rescoring reached 0.9190 with 213.38 MiB. TurboQuant 4-bit reached 0.8910 with 212.10 MiB. Binary 1-bit reached only 0.3215, despite using 198.76 MiB. The same general warning appeared here: a very small search copy can lose too much information.

Those numbers belong to this smaller, real-location workload. The baseline and quantized collections are approximate indexes built separately, so scalar finishing slightly above the approximate Float32 row is not proof that rounding makes exact search smarter. It means this particular indexed search returned more of the exact top-ten IDs.

Here the numbers are from the seeded fictional-scenario run, with 2× rescoring for quantized methods. Float32 cold is the unquantized approximate HNSW baseline. The full matrix is in Figure 1. Here are the observations I would remember.

Scalar int8 reached Recall@10 of 0.7690, versus 0.7385 for the Float32 cold approximate baseline, with 428.86 MiB of container memory versus 592.90 MiB. This does not prove that quantization improves exact accuracy. Separate approximate graph builds can return different candidates. The exact Float32 search, not the approximate Float32 row, is the 1.0 reference.

TurboQuant 4-bit matched that Float32 cold baseline’s 0.7385 recall while its container used 393.04 MiB. TurboQuant 2-bit used 377.24 MiB but recall fell to 0.6140. Binary 1-bit used 363.36 MiB, yet recall was only 0.2605. Product x32 used 366.32 MiB and reached 0.4810. Saving more memory did not automatically preserve the right results.

Median warm latency ranged from 5.830 to 10.131 ms across these rows. P95 was roughly 353 to 368 ms because filtered searches were much slower than unfiltered searches in this local setup. A single “speed” number would hide that split. With rescoring, cached bytes also changed because reading the original vectors warms pages. This is a trade-off, not a free memory discount.

Look at the two ends of each line. Float32 cold had 4.9 ms unfiltered P95, but 360.1 ms filtered P95. Scalar int8 with 2× rescoring had 4.0 ms and 364.3 ms. TurboQuant 4-bit had 4.4 ms and 369.5 ms. The pooled P95 near 350–370 ms is therefore a poor description of an ordinary unfiltered request. It mostly reflects the slow tail of this benchmark’s filtered half.

I cannot tell from this chart alone that filtering is the only cause. The filtered questions have different eligibility sets and may take different paths through the index. The honest conclusion is narrower: on this machine and query mix, the filtered tail needs separate investigation before anyone promises a production response time.

You can see the split in the saved CSV without rerunning Qdrant:

import csvwith open("results/scenarios-100k/summary.csv") as f:    rows = list(csv.DictReader(f))for row in rows:    if (row["configuration"] == "scalar_int8"            and row["mode"] == "rescore_2x"):        print(row["workload"], row["p95_ms"])

Yes. On the 100,000-point run, moving from no rescoring to 2× rescoring raised Recall@10 by 0.0595 for scalar, 0.0645 for TurboQuant 4-bit, 0.0980 for TurboQuant 2-bit, and 0.1175 for product quantization. Binary improved by 0.0700 but stayed low at 0.2605. Rescoring can rearrange the candidates it receives; it cannot bring back a good neighbour that never entered the candidate list.

The flat first step is useful. With no rescoring and with 1× rescoring, Recall@10 was the same for every quantized method in this run. Rechecking the same ten candidates can change their order, but it cannot add a neighbour that never became a candidate. At 2×, Qdrant has roughly twenty candidates to check for a top-ten answer. Scalar rises from 0.7095 to 0.7690; TurboQuant 4-bit rises from 0.6740 to 0.7385. Product x32 makes a larger jump, from 0.3635 to 0.4810, but still trails the two higher-recall methods.

At 4×, scalar stays at 0.7690, while product x32 rises again to 0.5515. That is why I would tune oversampling for a specific model instead of always choosing the largest number. More candidates ask Qdrant to read and score more original vectors. The extra work may be worthwhile if it recovers important results, but a flat curve gives me no reason to pay for it.

If I were responsible for this service, I would start with Scalar int8. It gave me the strongest Recall@10 in the 100,000-point run and reduced measured container memory from 592.90 MiB for Float32 cold to 428.86 MiB at 2× rescoring. My second experiment would be TurboQuant 4-bit because it matched the Float32-cold approximate baseline’s recall while using 393.04 MiB. I would not select Binary 1-bit just because “32×” looks impressive; I would have to explain why three quarters of the exact neighbours disappeared. Before deploying either choice, I would repeat the run on the target Linux machine with real driver queries and an agreed recall threshold. The result here gives me a sensible first configuration, not permission to skip production testing.

The 15,205-point run uses real charging-location descriptions, but it is not a live telemetry feed. The 100,000-point run tests scale with synthetic scenarios, and the two workloads have different texts and queries, so I cannot attribute every difference between them to point count alone. The one-million-point configuration exists, but no one-million-point measurement has been completed.

The tests ran on a Mac through Docker Desktop with a fixed container limit. The warm-search results do not tell us what a dedicated Linux server would do under heavy memory pressure or after a truly cold disk start. A production decision should repeat the experiment on the target hardware, with the target embedding model and real query mix.

The repository contains the data-preparation pipeline, seven Qdrant collection settings, a runnable notebook, tests, raw query outputs, memory snapshots, figures, and a report that recalculates Recall@10 from the returned IDs. The quick smoke run uses 4,096 artificial vectors and needs Docker but no dataset key or model download. For the real run, the code uses the fixed OpenChargeMap export and checksummed DfT archive. The OpenChargeMap API is linked for developers who want a current feed, but this measured run uses the export for reproducibility.

uv sync --lockeduv run ev-benchmark prepare --config configs/smoke.jsonuv run ev-benchmark run --config configs/smoke.json --output runs/smoke

The path runs/smoke/report.md refers to a generated Markdown file inside the cloned repository, not to a public webpage. After the two smoke commands finish, open that file in GitHub, VS Code, or any text editor. It summarizes the 4,096-point artificial fixture and confirms that collection creation, queries, memory capture, and report generation work end to end. It is a quick installation check rather than the evidence used for the article’s headline numbers. The measured reports are results/real-gb/report.md and results/scenarios-100k/report.md, and both are also linked in the Sources section below. The public repository is at https://github.com/Pranshu640/qdrant-ev-memory-benchmark.

I started with a simple worry: as an EV search service grows, will its vectors force every byte into expensive RAM? This experiment gave me a more useful answer than a single compression ratio. The small structures used to find candidates can stay fast and close to the CPU, while the full Float32 vectors wait in cold mmap storage until rescoring needs a shortlist. In this workload, Scalar int8 was the safest first choice; TurboQuant 4-bit was the more aggressive option I would test next. Binary showed why memory efficiency cannot be judged without recall. The broader lesson is the one I would carry into a million-point run: choose the memory layout and the quantizer together, then measure the bytes saved against the neighbours lost. A “32×” label is only the beginning of that decision.

Project repository and notebook: https://github.com/Pranshu640/qdrant-ev-memory-benchmark

Measured real-location report: https://github.com/Pranshu640/qdrant-ev-memory-benchmark/blob/main/results/real-gb/report.md

Measured 100,000-scenario report: https://github.com/Pranshu640/qdrant-ev-memory-benchmark/blob/main/results/scenarios-100k/report.md

Qdrant quantization guide: https://qdrant.tech/documentation/manage-data/quantization/

Qdrant memory tiers and legacy on_disk mapping: https://qdrant.tech/documentation/ops-configuration/memory-tiers/

Qdrant TurboQuant release notes: https://github.com/qdrant/qdrant/releases/tag/v1.18.0

Qdrant on-disk storage guide: https://qdrant.tech/documentation/manage-data/storage/

OpenChargeMap API: https://www.openchargemap.org/develop/api

OpenChargeMap export used here: https://github.com/openchargemap/ocm-export

DfT road traffic catalogue: https://www.data.gov.uk/dataset/208c0e7b-353f-4e2d-8b7a-1a7118467acc/gb-road-traffic-counts

DfT download page: https://roadtraffic.dft.gov.uk/downloads

Google Research explanation of TurboQuant: https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/

I Tested Qdrant Quantization and mmap on EV Charger Search was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @qdrant 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-tested-qdrant-quan…] indexed:0 read:21min 2026-09-29 · —