The finding that should change how you think about this work: newer, larger models did not produce more secure code than their predecessors.

Key takeaways

  • Veracode's GenAI Code Security Report tested more than 100 large language models across Java, JavaScript, Python and C#, and found AI-generated code contained 2.74 times more.
  • 124 live listings in this category publish a rate, at a median top-of-range of $100 an hour and a ceiling of $300.
  • The work is remote contract work, asynchronous, with no set hours and no guaranteed volume.
  • Applications screen on a short skills assessment rather than a resume or interview.

What was reported

The finding

Veracode's GenAI Code Security Report tested more than 100 large language models across Java, JavaScript, Python and C#, and found AI-generated code contained 2.74 times more vulnerabilities than human-written code. 45% of AI-generated samples introduced OWASP Top 10 vulnerabilities. Java performed worst with a 72% security failure rate, and cross-site scripting had an 86% failure rate. The report notes this has remained largely unchanged even as models improved at producing syntactically correct code.

Veracode ran more than 100 models across four languages. AI-generated code carried 2.74 times the vulnerabilities of human-written code, 45% of samples introduced an OWASP Top 10 issue, and Java failed security checks 72% of the time. Models got much better at writing code that compiles and no better at writing code that is safe.

What the listings pay

That is not a gap that closes with the next release, and it is the reason code evaluation is a durable category rather than a temporary one. Across 124 live engineering and code listings on our board that publish a rate, the median top-of-range is $100 an hour, reaching $300.

#RoleAdvertised ratePlatform
1CUDA Engineering Expert$300 to $300 an hourMercor
2Cybersecurity Research Expert, Offensive Security & Vulnerability Research$200 to $250 an hourMercor
3Machine Learning Engineer Talent Network$70 to $250 an hourMercor
4Legacy Codebase Migration Expert$200 to $200 an hourMercor
5UK-Based Data Engineering Experts$140 to $200 an hourMercor
6AI Software Engineering Domain Expert$100 to $200 an hourmicro1
7AI/ML Engineer$70 to $200 an hourmicro1
8Engineering / Platform Professionals$80 to $160 an hourMercor
9Open Source Applied Engineer Talent Network$100 to $150 an hourMercor
10Backend Engineer Talent Network$70 to $150 an hourMercor

Source: 124 live listings on this board that publish a rate, read directly from each posting on 2026-09-06. Listings without a published rate are excluded rather than estimated.

AI Code Has 2.74x More Vulnerabilities. Someone Has to Find Them

How this compares across the board

A rate only means something next to the alternatives. This is every category we track with at least five listings publishing a rate, ranked by median top-of-range, so you can see where this work sits rather than taking a single number on trust.

CategoryListingsMedian lowMedian topHighest
Legal95$100$140$400
Medical68$77$120$400
Consulting45$80$120$280
Finance94$80$110$280
Engineering114$70$100$300
Research/PhD132$70$90$280
Writing36$40$80$280
Bilingual78$44$52$120
Annotation25$12$24$120

Same source and date as above. Categories are matched on listing title, so a role can appear in more than one.

What it means for you

The skill being paid for is specific and worth naming: reading code that compiles, looks correct, and is quietly unsafe. That is a different muscle from writing code, and it is the one in demand. See code review and debugging roles.

Who should apply

Two checks before you spend time on an application. Confirm the role accepts applicants from your country with the eligibility checker, since a meaningful share of listings carry location requirements. Then run the advertised rate through the take-home calculator, because this is contract work and the headline figure is before self-employment tax.

Applications complete on the hiring platform and usually take a few minutes, with a short skills assessment in place of an interview. Fill in every credential, language and professional background field on your profile. Those are what route you to the better paid listings, and most applicants leave them blank.

Frequently asked questions

How much less secure is AI-generated code?

Veracode found 2.74 times more vulnerabilities than human-written code, with 45% of samples introducing an OWASP Top 10 vulnerability.

Which languages performed worst?

Java was worst at a 72% security failure rate. Cross-site scripting was the worst individual category at an 86% failure rate.

Will newer models fix this?

The report found security performance largely unchanged as models improved, with newer and larger models not producing significantly more secure code.

What does code evaluation work pay?

Across 124 live engineering and code listings publishing a rate, the median top-of-range is $100 an hour, reaching $300.

Do I need a security background?

Helpful but not usually required. Most listings want strong engineers who read unfamiliar code carefully; dedicated security roles do ask for the background.

Sources

  1. Veracode, AI-generated code: a double-edged sword for developers
  2. Cloud Security Alliance, AI-generated code vulnerability surge

See every live role

The full board updates several times a week, with the advertised rate on each listing and closed roles removed.

Browse all AI jobs