Remote | Open Source Software Engineer — $60–$80/hour

24-Mag

  • New York, New York
  • 6 days ago
  • Remote

    Highlights

    Selected engineers will audit repository-level assignments, reference patches, test harnesses, containerised environments, and evaluation logic while identifying technical flaws, unintended shortcuts, and weaknesses in task design. We are sharing a specialised part-time consulting opportunity for experienced Software Engineers with hands-on open-source contribution or maintainer experience and strong expertise in repository-level code review, testing, debugging, and software quality evaluation.

    Numbers & Facts

    LocationNew York, New York (
    Remote
    )
    Website4-mag.com/privacy-policy

    Description

    We are sharing a specialised part-time consulting opportunity for experienced Software Engineers with hands-on open-source contribution or maintainer experience and strong expertise in repository-level code review, testing, debugging, and software quality evaluation.

    This role focuses on reviewing software-engineering benchmark tasks for correctness, reproducibility, and grading integrity. Selected engineers will audit repository-level assignments, reference patches, test harnesses, containerised environments, and evaluation logic while identifying technical flaws, unintended shortcuts, and weaknesses in task design.

    Key Responsibilities

    Repository-Level Code Review

    • Review software engineering tasks built around real code repositories
    • Assess whether task requirements are technically clear, complete, and reproducible
    • Evaluate repository state, dependencies, configuration, and expected behaviour
    • Identify ambiguities or implementation issues that could affect task validity
    • Apply practical engineering judgement to realistic codebase-level problems

    Reference Patch Auditing

    • Review reference patches for correctness and completeness
    • Determine whether proposed solutions appropriately address the underlying software issue
    • Identify unintended behavioural changes, incomplete fixes, or unsupported assumptions
    • Compare reference implementations against task requirements and expected outcomes
    • Assess whether alternative valid implementations are treated fairly

    Test Harness & Grading Review

    • Audit test runners and automated evaluation logic
    • Assess whether tests accurately measure the intended behaviour
    • Identify missing coverage, brittle assertions, or grading inconsistencies
    • Verify that evaluation criteria appropriately distinguish correct from incorrect solutions
    • Review benchmark tasks for reliable and repeatable scoring

    Reproducibility & Environment Validation

    • Evaluate whether tasks can be reproduced consistently across clean environments
    • Review dependency installation, build processes, configuration, and runtime requirements
    • Assess Docker-based isolation and containerised execution
    • Identify environmental dependencies or hidden assumptions affecting reproducibility
    • Verify that tasks execute reliably under their intended setup

    Benchmark Integrity

    • Identify potential answer leakage, unintended shortcuts, or reward-hacking opportunities
    • Evaluate whether benchmark structure exposes information that makes tasks artificially easy
    • Review task and grading design for loopholes or exploitable behaviours
    • Assess whether successful completion genuinely demonstrates the intended engineering capability
    • Recommend improvements where benchmark integrity is compromised

    Software Testing & Debugging

    • Investigate failing or inconsistent benchmark tasks
    • Review stack traces, logs, test failures, and repository behaviour
    • Identify root causes of technical issues
    • Distinguish task defects from legitimate implementation failures
    • Assess whether debugging and validation processes follow sound engineering practices

    Open-Source Engineering

    • Apply experience from contributing to or maintaining open-source software
    • Evaluate repository conventions, contribution patterns, and realistic development workflows
    • Review patches with the perspective of an experienced contributor or maintainer
    • Assess whether proposed changes would meet reasonable code-review expectations
    • Apply practical judgement derived from real-world pull request and repository experience

    Multi-Language Code Evaluation

    • Review software written in Python
    • Evaluate tasks involving at least one additional ecosystem such as Java, Go, TypeScript, or C++
    • Assess code structure, tests, implementation choices, and repository conventions across languages
    • Identify language-specific implementation or testing issues
    • Apply consistent engineering standards across different technology stacks

    Rubric-Based Evaluation

    • Assess benchmark tasks against structured technical criteria
    • Provide clear written explanations supporting evaluation decisions
    • Reference specific code, tests, patches, or execution behaviour when identifying issues
    • Apply grading standards consistently across assignments
    • Distinguish substantive benchmark defects from minor implementation differences

    Ideal Profile

    • 3+ years of professional software engineering experience
    • Demonstrated open-source contribution or maintainer experience, such as merged pull requests, committer responsibilities, or maintainer roles
    • Strong ability to review repository-level software changes
    • Experience auditing reference patches, test runners, and automated test suites
    • Comfortable evaluating Docker-based isolation and reproducible development environments
    • Strong understanding of software testing, debugging, and code-review practices
    • Ability to identify answer leakage, reward hacking, or other benchmark-integrity issues
    • Strong proficiency in Python
    • Professional fluency in at least one additional language such as Java, Go, TypeScript, or C++
    • Familiarity with SWE-Bench Verified or similar repository-level software engineering benchmarks is preferred
    • Maintainer or contributor history with established Python open-source projects is highly valued
    • Previous code-review, software evaluation, or task-grading experience is advantageous
    • Strong written communication and ability to provide precise technical feedback

    Engagement Details

    • Part-time independent contractor engagement
    • Fully remote within the United States
    • Flexible scheduling based on project requirements
    • Compensation: $60–$80/hour
    • Work focuses on repository-level software evaluation, reference-patch review, testing, reproducibility, benchmark integrity, and technical quality assessment
    • Projects may be extended, shortened, or concluded based on project needs and performance
    • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
    • H1-B and STEM OPT support is unavailable for this engagement

    About the Platform

    This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

    By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

    Similar Jobs