Skip to content

Proposed faster approach for --create imports #2504

Description

@leijurv

Hello,

I would like to propose contributing an approach that significantly speeds up the initial planet import on "ordinary" hardware by avoiding the node location random reads. In particular, I have a prototype implementation that sped up processing time from 71 hours to under 2 hours, and the total time including post-processing from 77 hours to under 9 hours, on a box with 8 CPU cores and 32GB RAM. (I'm defining "ordinary hardware" to mean "node locations don't fit in RAM, so the current approach uses the flat file on disk").

Phase Time before Time after
Scatter n/a (new) 0h 7m
Nodes 0h 46m 0h 19m
Ways 45h 41m 1h 7m
Relations 24h 33m 0h 8m
Processing total
(sum so far)
71h 0m 1h 43m
Post-processing 5h 59m 6h 54m
Total 76h 59m 8h 37m

Brief outline:

  • In a new phase (scatter), we read all the PBF ways, and save to disk all the (node_id, way_id) pairs.
  • Hybrid RAM+disk sort those by node_id
  • Join the (node_id, way_id) against the PBF (node_id, lat, lng). Both of these are now in ascending order by node_id (the PBF has it sorted in the first place), so we can do it streaming, zippering them together. Also at this time we run the Lua process_node in parallel. We now have entries of (way_id, node_id, lat, lng).
  • Hybrid RAM+disk sort those by way_id (back to the original order).
  • Read these, in way order, and again join it against the PBF for the ways. Within each way, we sort (one last time) to get the (lat, lng) list in the way's order. Lua process_way then happens.

Relations happen similarly (node members scattered in scatter, way members scattered in ways, gathered in relations).

Middle and Lua still see all nodes, then all ways, then all relations. The tables all end up with the same contents and indexes and clustering.

Details (click to expand)
  • Numbers from a m6g.2xlarge instance, with the boot disk increased to 2.5tb (EBS gp3) but nothing else changed. Full planet from Sept 14. OSM Carto flex style with --slim. Postgres 18.
  • The average CPU usage before (usr+sys) was 3.4% (only one core and lots of iowait). After the average was 82.7%. During processing. During post-processing unchanged at ~20%.
  • It writes a lot more to disk as temp files, ~250gb, but it's all sequential, and it does not change the peak disk usage since we delete it as we go.
  • Peak disk usage during the processing stage is 793gb before, 844gb after. The eventual peak is unchanged later on, 1346gb during clustering.
  • EBS stats. Processing stage only. Beforehand 22.5TB read across 362mil requests average 62kb each (of which only a tiny fraction was used). Afterward 0.69TB read across 5mil requests average 138kb each (of which a much larger fraction was used).
  • The CLUSTER, CREATE INDEX, etc post-processing got about an hour longer because I dropped the primary key from planet_osm_ways for faster loading, then added it back later.
  • Builds on top of Use binary COPY for flex and middle tables #2503
  • Lua is now run multithreaded, each worker thread having its own Lua state and COPY connection.
  • Currently only supports --create, flex, one input file, no bbox, no two-stage processing.
  • The "Hybrid RAM+disk sort" means the writer makes a few dozen bucket files, partitioned by the most significant bits of the sorting key. pseudocode: for way in ways: for node in way: bucket_files[node.id >> 28].buffered_write((node.id, way.id)) in the initial write (buffered in that we only actually write to disk every megabyte-ish). Then on the reader side, each bucket is now small enough to comfortably sort in RAM, and each bucket is all entirely lesser than the next bucket, so this results in a fully sorted stream.

Prototype first draft code is here: leijurv@2b6d018 The main file is here (it's not crazy long; 1.3k lines).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions