You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
I would like to propose contributing an approach that significantly speeds up the initial planet import on "ordinary" hardware by avoiding the node location random reads. In particular, I have a prototype implementation that sped up processing time from 71 hours to under 2 hours, and the total time including post-processing from 77 hours to under 9 hours, on a box with 8 CPU cores and 32GB RAM. (I'm defining "ordinary hardware" to mean "node locations don't fit in RAM, so the current approach uses the flat file on disk").
Phase
Time before
Time after
Scatter
n/a (new)
0h 7m
Nodes
0h 46m
0h 19m
Ways
45h 41m
1h 7m
Relations
24h 33m
0h 8m
Processing total (sum so far)
71h 0m
1h 43m
Post-processing
5h 59m
6h 54m
Total
76h 59m
8h 37m
Brief outline:
In a new phase (scatter), we read all the PBF ways, and save to disk all the (node_id, way_id) pairs.
Hybrid RAM+disk sort those by node_id
Join the (node_id, way_id) against the PBF (node_id, lat, lng). Both of these are now in ascending order by node_id (the PBF has it sorted in the first place), so we can do it streaming, zippering them together. Also at this time we run the Lua process_node in parallel. We now have entries of (way_id, node_id, lat, lng).
Hybrid RAM+disk sort those by way_id (back to the original order).
Read these, in way order, and again join it against the PBF for the ways. Within each way, we sort (one last time) to get the (lat, lng) list in the way's order. Lua process_way then happens.
Relations happen similarly (node members scattered in scatter, way members scattered in ways, gathered in relations).
Middle and Lua still see all nodes, then all ways, then all relations. The tables all end up with the same contents and indexes and clustering.
Details (click to expand)
Numbers from a m6g.2xlarge instance, with the boot disk increased to 2.5tb (EBS gp3) but nothing else changed. Full planet from Sept 14. OSM Carto flex style with --slim. Postgres 18.
The average CPU usage before (usr+sys) was 3.4% (only one core and lots of iowait). After the average was 82.7%. During processing. During post-processing unchanged at ~20%.
It writes a lot more to disk as temp files, ~250gb, but it's all sequential, and it does not change the peak disk usage since we delete it as we go.
Peak disk usage during the processing stage is 793gb before, 844gb after. The eventual peak is unchanged later on, 1346gb during clustering.
EBS stats. Processing stage only. Beforehand 22.5TB read across 362mil requests average 62kb each (of which only a tiny fraction was used). Afterward 0.69TB read across 5mil requests average 138kb each (of which a much larger fraction was used).
The CLUSTER, CREATE INDEX, etc post-processing got about an hour longer because I dropped the primary key from planet_osm_ways for faster loading, then added it back later.
Lua is now run multithreaded, each worker thread having its own Lua state and COPY connection.
Currently only supports --create, flex, one input file, no bbox, no two-stage processing.
The "Hybrid RAM+disk sort" means the writer makes a few dozen bucket files, partitioned by the most significant bits of the sorting key. pseudocode: for way in ways: for node in way: bucket_files[node.id >> 28].buffered_write((node.id, way.id)) in the initial write (buffered in that we only actually write to disk every megabyte-ish). Then on the reader side, each bucket is now small enough to comfortably sort in RAM, and each bucket is all entirely lesser than the next bucket, so this results in a fully sorted stream.
Prototype first draft code is here: leijurv@2b6d018 The main file is here (it's not crazy long; 1.3k lines).
Hello,
I would like to propose contributing an approach that significantly speeds up the initial planet import on "ordinary" hardware by avoiding the node location random reads. In particular, I have a prototype implementation that sped up processing time from 71 hours to under 2 hours, and the total time including post-processing from 77 hours to under 9 hours, on a box with 8 CPU cores and 32GB RAM. (I'm defining "ordinary hardware" to mean "node locations don't fit in RAM, so the current approach uses the flat file on disk").
(sum so far)
Brief outline:
scatter), we read all the PBF ways, and save to disk all the(node_id, way_id)pairs.node_id(node_id, way_id)against the PBF(node_id, lat, lng). Both of these are now in ascending order bynode_id(the PBF has it sorted in the first place), so we can do it streaming, zippering them together. Also at this time we run the Luaprocess_nodein parallel. We now have entries of(way_id, node_id, lat, lng).way_id(back to the original order).(lat, lng)list in the way's order. Luaprocess_waythen happens.Relations happen similarly (node members scattered in
scatter, way members scattered inways, gathered inrelations).Middle and Lua still see all nodes, then all ways, then all relations. The tables all end up with the same contents and indexes and clustering.
Details (click to expand)
m6g.2xlargeinstance, with the boot disk increased to 2.5tb (EBS gp3) but nothing else changed. Full planet from Sept 14. OSM Carto flex style with--slim. Postgres 18.CLUSTER,CREATE INDEX, etc post-processing got about an hour longer because I dropped the primary key fromplanet_osm_waysfor faster loading, then added it back later.--create, flex, one input file, no bbox, no two-stage processing.for way in ways: for node in way: bucket_files[node.id >> 28].buffered_write((node.id, way.id))in the initial write (buffered in that we only actually write to disk every megabyte-ish). Then on the reader side, each bucket is now small enough to comfortably sort in RAM, and each bucket is all entirely lesser than the next bucket, so this results in a fully sorted stream.Prototype first draft code is here: leijurv@2b6d018 The main file is here (it's not crazy long; 1.3k lines).