# Multi-threaded decompression

**URL:** <https://forum.kx.com/t/multi-threaded-decompression/10941>\
**Category:** Community Support\
**Tags:** kdb-and-q\
**Created:** [April 9, 2016, 4:18am UTC](https://forum.kx.com/t/multi-threaded-decompression/10941 "2016-04-09T04:18:00Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Victor\_Wong1](https://avatars.discourse-cdn.com/v4/letter/v/b77776/32.png) [@Victor\_Wong1](https://forum.kx.com/u/Victor_Wong1)\
**Post date:** [April 9, 2016, 4:18am UTC](https://forum.kx.com/t/multi-threaded-decompression/10941/1 "2016-04-09T04:18:00Z")

</div>

Hi,

I am wondering why multi-threaded decompression (using peach) appears to be slower than using a single thread. Is there some sort of IPC between threads that’s causing extra overhead? What is the right way to decompress in parallel?

Thanks in advance,

Victor

c:\q\w32\>q -s 3

KDB+ 3.3 2016.03.14 Copyright (C) 1993-2016 Kx Systems

Welcome to kdb+ 32bit edition

For support please see [http://groups.google.com/d/forum/personal-kdbplus](http://groups.google.com/d/forum/personal-kdbplus)

Tutorials can be found at [http://code.kx.com/wiki/Tutorials](http://code.kx.com/wiki/Tutorials)

To exit, type \

To remove this startup msg, edit q.q

q).z.zd:(16;2;6)

q)(@[`:c:/db/t;;:;].‘) ((`sym;10000000?($[`]’)“abcde”);(`time;.z.Z - 10000000?1000f);(`price;100 - (10000000?20f) - 10))

`:c:/db/t`:c:/db/t`:c:/db/t

q)\t @[`:c:/db/t;] each `time`sym`price

468

q)\t @[`:c:/db/t;] peach `time`sym`price

982

---

<div class="post-metadata">

**Author:** ![Victor\_Wong1](https://avatars.discourse-cdn.com/v4/letter/v/b77776/32.png) [@Victor\_Wong1](https://forum.kx.com/u/Victor_Wong1)\
**Post date:** [April 10, 2016, 10:20am UTC](https://forum.kx.com/t/multi-threaded-decompression/10941/2 "2016-04-10T10:20:00Z")

</div>

I realize the motivation may not be clear from the example above, so I put together some use cases below. &nbsp;From testing, it looks like select doesn’t do compressed column scans - for filtering or retrieval - in parallel, and one can outperform select by using peach. &nbsp;However, it doesn’t generalize when one tries to parallelize both. &nbsp;I imagine there is probably some setting/library I am not aware of to optimize queries on compressed tables given how long compression has been part of kdb, so if anyone has any suggestions, I’d really appreciate it.

`c:\q\w32>q -s 3KDB+ 3.3 2016.03.14 Copyright (C) 1993-2016 Kx Systemsw32/ 4()core 4095MB NONEXPIREq).z.zd:(16;2;6)q)`:c:/db/t/ set .Q.en[`:c:/db/] ([]sym:10000000?($[`]')“abcde”;time:.z.Z - 10000000?1000f;price:100 - (10000000?20f) - 10)`:c:/db2/t/q)\l c:/db`

Retrieving columns in parallel using peach outperforms simple select.

`q)\t select from t where sym in `a`b`c1560q)\t {[t;s] flip c!{x[z] y}[t;exec i from t where sym in s] peach c:cols t}[t;`a`b`c]1170`

Filtering rows across multiple columns using peach also outperforms standard select.

`q)\t exec i from t where (sym in `a`b`c) and (time \> 2015.01.01) or price \> 1001669q)\t {x inter y union z} . {eval parse"exec i from t where ",x} peach ("sym in `a`b`c";"time > 2015.01.01";"price > 100")1357`

However, it is actually slower when you try to do both in parallel.

`q)\t select from t where (sym in `a`b`c) and (time \> 2015.01.01) or price \> 1001747q)\t {[t;s] flip c!{x[z] y}[t;s] peach c:cols t}[t] {x inter y union z} . {eval parse"exec i from t where ",x} peach ("sym in `a`b`c";"time > 2015.01.01";"price > 100")2090`

---

<div class="post-metadata">

**Author:** ![charlie1](https://avatars.discourse-cdn.com/v4/letter/c/ec9cab/32.png) [@charlie1](https://forum.kx.com/u/charlie1)\
**Post date:** [April 11, 2016, 8:20am UTC](https://forum.kx.com/t/multi-threaded-decompression/10941/3 "2016-04-11T08:20:00Z")

</div>

kdb+ uses thread local heaps, and uses serialization when passing data back from slave threads to the main thread during peach. In general, the problems suitable for peach are those that incur a high computation cost returning small data.

kdb+ has built in support for using multiple threads for the … in queries such as

select … by s from t where s in S

where s has a p or g attr. In addition it multithreads queries across partitions.

If you’re not doing any aggregation, or you don’t have a g or p attr on s, then it’s possible you’ll find manually crafted explicit routes which perform better.

---

<div class="post-metadata">

**Author:** ![Victor\_Wong1](https://avatars.discourse-cdn.com/v4/letter/v/b77776/32.png) [@Victor\_Wong1](https://forum.kx.com/u/Victor_Wong1)\
**Post date:** [April 18, 2016, 2:21am UTC](https://forum.kx.com/t/multi-threaded-decompression/10941/4 "2016-04-18T02:21:00Z")

</div>

Thanks, Charles. &nbsp;So if I understand you correctly, it uses multiple threads when the query spans multiple partitions, or performs aggregations over partitioned or indexed columns, but does not differentiate between compressed and uncompressed columns?
