Skip to content

Run new benchmarks and document costs #70

Description

@ehoelzl

Supersedes mlbench/mlbench-core#82. We can now also use PyTorch 1.7.0

  • CIFAR10, ResNet20, All Reduce, 1 to 16 workers
  • CIFAR10, ResNet20, DDP, 1 to 16 workers
  • Wikitext2, LSTM, All Reduce, 1 to 16 (32 ?) workers
  • Wikitext2, LSTM, DDP, 1 to 16 (32 ?) workers
  • WMT16, LSTM, All Reduce, 1 to 32 workers
  • WMT16, LSTM, DDP, 1 to 32 workers
  • WMT17, Transformer, All Reduce, 1 to 32 workers
  • WMT17, Transformer, DDP, 1 to 32 workers

Activity

  1. ehoelzl commented on Dec 4, 2020

    @ehoelzl
    ContributorAuthor

    @martinjaggi 's comment
    "
    there are many use-cases which need to be run

    • the official task results (light and full goals need to run)
    • scaling as part of official results
    • comparing different hardware such as gcloud GPUs K80,P100,P4,T4,V100 etc
    • GPU vs CPU (though official results will be GPU only)
    • comparing backends (probably no scaling the # workers needed)
    • comparing against pytorch DDP or powerSGD, or adam (again this is different than the official results so has different
    • regression testing compared to old versions of our code, and changes in pytorch, TF say
      depending on what's needed for the scientific report (or blog post), that's what we'll run. ideally the code would be easy to adapt
      "
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions