5 comments

  • happyPersonR a day ago ago

    i'm glad they had a skeptic

    might be overkill/flawed , but maybe something that checks to see using a VLM that the GIF's that are generated/modified etc using the old lib and the new one look similar ish at least?

    if you have google's tpu's ... could probably be done

    also an agentic flow that generate's gif's that might break the C version or actually cause pointer type exploits and see if similar approaches work or don't in the rust version?

    maybe even an agentic red team that tries to do bad things to this through gif's to the library and see what it does?

    just thoughts... need coffee haha :)

    • kgen 20 hours ago ago

      Yes, from the article they basically did all those things

      • happyPersonR 5 hours ago ago

        some but not all:

        they did " Mass-Scale Regression Testing: Leveraging our internal data processing infrastructure, we validated the rewrite against a dataset of over 30 million GIFs. This ensured that the new implementation produced results identical to the original version across a large set of real-world inputs."

        didn't say they did anything with a VLM to check the output vs the original library

        "Differential Fuzzing: We implemented a differential fuzzer that exercised the original C and new Rust implementations side-by-side. In over six days of continuous execution, the fuzzer performed over 200 million iterations without identifying any logic deviations."

        this just checks parity/equivalence.. i'm talking about new fuzzing just for the new rust library as a harness. This isn't that.

        'Adversarial AI Analysis: We utilized specialized LLM prompts to perform "adversarial reviews," asking models to identify subtle behavioral differences between the two codebases that might escape traditional testing.'

        this is just looking at the code...

        i'm sure they probably did all the other things google's internal standards required.. but tbh, i would hope google are updating and enhancing those for the agentic harness era. Now that we can we should :)

  • ramon156 a day ago ago

    it feels yucky to read LLM generated words on a Google blog, but i guess it's not too distracting. cool write-up though

    • pjmlp 13 hours ago ago

      Some people apparently are not coding any longer, Claude and friends do their work, so why wouldn't they write their posts as well?

      If it is to replace the human, it has to be in all places.