New CRAM method offers giant boost to compressed memory reads

(tomshardware.com)

58 points | by danny00 2 days ago ago

26 comments

  • Kevin_Flynn 2 days ago ago

    Ram Doubler, for Mac and Win 3.x, 1996

    https://winworldpc.com/product/connectix-ram-double/windows-...

    Nice to see its back, with better performance.

    History repeated, with refinement.

  • eglintondust a day ago ago

    So can I finally download more RAM?

    • smallmancontrov a day ago ago

      https://downloadmoreram.com/

      Whoa, new options! I'm downloading the Quantum RAM with Haptic Feedback.

    • skavi a day ago ago

      probably not to your current computer. this works on supporting memory types that do compression in hardware.

    • swed420 a day ago ago

      You wouldn't download a car

      • a day ago ago
        [deleted]
  • cyberclimb a day ago ago

    the recording of the presentation has got to be on this YT channel but I only scrubbed through 2 videos (8h each) and it's not easy to find

    https://www.youtube.com/@LinuxPlumbersConference

    not sure where it falls in the schedule here either https://lpc.events/event/20/timetable/#all

  • vfosnar a day ago ago
    • dang a day ago ago

      Added above. Thanks!

    • x0z a day ago ago

      Thank you <3

  • cwillu a day ago ago

    The article appears to misunderstand the point of the project, and mostly just describes what zswap already does (writing compressed pages back into ram), rather than talking about cram's hardware-offloaded compression that allows cacheline-level access rather than page-level. (And they appear to be aware that they don't really understand it: “My explanation of CRAM might not be completely correct”)

    The phoronix article is better, and I say that as someone who usually detests the quality of phoronix's technical writing.

    • tancop a day ago ago

      It's not really about hardware offload, the biggest problem with zswap is that it's swap. The rest of the kernel treats it like a fast SSD (which is still incredibly slow compared to RAM) instead of slower memory that needs a bit of special handling on writes.

      NUMA maps a lot closer to what compressed RAM actually is. The subsystem is more aware of CRAMs specifics so it can make better decisions about where to put allocations and everything gets faster. And it's less overhead because swap is not really optimized for frequent direct access but for NUMA it's the most basic function.

      • ahartmetz a day ago ago

        Whoa, building it on top of the NUMA logic is pretty clever - also, finally all that complicated code does something useful on normal computers!

    • ragall 6 hours ago ago

      The article was written by an LLM. No understanding was involved, no brain cells harmed in the process.

  • madduci a day ago ago

    Finally we can run frontier models locally

    • sroussey a day ago ago

      We brute force AI models right now because a) we don’t know better, and b) it’s premature optimization.

      I beg to differ on point b, but no one is delaying their next model just so they can concentrate on optimization.

      It’s coming though.

      One example: https://siliconangle.com/2026/07/28/ai-model-compression-sta...

      Another is separating the the intelligence part of the model from the known facts part of the model (which can be better compressed)

      • djmips 19 hours ago ago

        The Chinese breakthroughs are often around optimization. And I gather it's because they are constrained and very bright so it's a natural progression.

        • ranger_danger 13 hours ago ago

          IMO Restrictions are exactly what sparks creativity in people... whereas too many options leads to choice paralysis, which is what I fear is a big problem in the FOSS world.

          Don't get me wrong... I think having choices is still good, I just think that having too many choices is not equally as good, or necessarily better.

    • throawayonthe a day ago ago

      one would hope model weights are already high entropy enough and would not be improved by a general-purpose memory compression algo?

    • tintor a day ago ago

      Compression doesn't really work for model weights.

      Model quantization and model distillation are two techniques to reduce model size.

  • yonatan8070 21 hours ago ago

    Is the hardware support this needs common? Can it be used on regular x86-64 desktops/laptops or is it for specialized hardware?