Semgrep vs GitHub CodeQL: Pattern Rules or Semantic Queries
Semgrep matches code shapes without a build; CodeQL queries a semantic database of your program. How that choice plays out in ownership, noise and depth.
At a glance
| Fact | Semgrep Semgrep, Inc. | GitHub CodeQL GitHub |
|---|---|---|
| Best for | Writing and enforcing custom code patterns across polyglot repositories | Writing precise custom queries to find vulnerability variants at scale |
| License | Open source and commercial | Open source and commercial |
| Maturity | Established | Established |
| Deployment 2 in common |
|
|
| Languages 10 in common |
|
|
| Integrations 3 in common |
|
|
| Profile | Full Semgrep profile Some details still being confirmed | Full GitHub CodeQL profile Some details still being confirmed |
From the structured catalog records. Highlighted entries are shared by both tools. No scores or rankings: see the editorial policy.
Contents
These two land on the same shortlist for a predictable reason. Both are open at the core, both let your own engineers write detection logic instead of waiting on a vendor, and both treat the shipped rule set as a starting point rather than the product. Teams that have outgrown a scanner they cannot tune, or that just had an incident and want to make sure the same mistake is not sitting in every other service, usually end up comparing exactly this pair.
The difference the rest of this article turns on is what each tool thinks your code is. Semgrep sees a syntax tree and asks whether a fragment of it has a particular shape. CodeQL sees a database of your program, with types, control flow and a call graph, and asks logical questions about it. Nearly every practical consequence, from onboarding speed to what finds its way into the triage queue, follows from that one design decision.
The short answer #
Pick Semgrep if you run many languages, want scanning on every repository quickly, and have a security engineer who will encode your internal conventions as rules. Pick CodeQL if your code lives on GitHub, your builds are healthy, and someone on the team will learn QL well enough to model your frameworks and hunt variants. Running both is reasonable in a larger program: Semgrep as the fast, broad pull request gate for conventions and banned APIs, CodeQL for deep data flow analysis on the services where an injection bug would hurt most.
How they differ #
Detection approach #
A Semgrep rule looks like the code it catches. You write a call with metavariables in place of the parts that vary, add operators for containment and exclusion, and the engine matches it against the parsed tree, so formatting and comments do not break it. Taint mode adds sources, sinks and sanitizers for cases where a shape match is not enough. The boundary matters: in the open source engine that taint analysis stays within a file, and cross file and cross function tracking belong to the commercial engine.
CodeQL extracts the program into relational tables and runs queries in QL, a declarative logic language with classes, predicates and recursion. Its security packs sit on shared taint tracking libraries that follow data across functions and files, and the full source to sink path comes back with each result. You can ask for every flow from a request parameter to a deserialization call that does not pass through your validator, and the answer respects how the program actually connects.
This axis favors CodeQL for anyone whose main concern is interprocedural data flow, and Semgrep for anyone whose risk lives in recognizable code shapes.
Setup and ownership #
Semgrep needs source and nothing else. No build, no dependency installation, no extractor to configure, which is why it spreads across a polyglot estate in days. Ownership mostly means curating rules: deciding which registry rules to trust, writing your own, and keeping the enabled set tight.
CodeQL on interpreted languages is similarly easy, but compiled languages need a working build for extraction. That puts the scanner downstream of your build health, and a broken extraction quietly becomes a coverage gap. On GitHub the Actions workflow and security tab make this smooth. Elsewhere you run and maintain the CLI yourself and ship SARIF into whatever consumes it. Also confirm the license terms that apply to scanning private code in your setup before planning a rollout.
This axis favors Semgrep for teams with limited platform capacity or messy builds, and CodeQL for GitHub native organizations with reliable CI.
Signal quality at scale #
Semgrep's noise usually comes from two places: registry rules of uneven precision enabled wholesale, and syntactic rules firing where safety depends on context the rule cannot see, such as validation that happened several calls earlier. Diff aware scanning, which reports only findings the change introduced, keeps this manageable on legacy code, and disciplined curation keeps it low.
CodeQL's semantic model lets a query recognize sanitizers and follow data through wrappers, which removes a class of false positives that plague shape matching. The trade is that when its model of your framework is incomplete, it misses flows instead of flagging them, and nothing in the output tells you that happened. Adding your custom request objects and sanitizers to the model is what turns a decent default into a precise one.
This axis favors CodeQL once someone has modeled your frameworks, and Semgrep when you want predictable, readable findings from day one.
Extensibility and variant hunting #
Both tools are built to be extended, but they reward different people. Most developers can read a Semgrep rule during code review and write one in an afternoon, so rule authorship can spread beyond the security team. That matters when your goal is enforcing conventions like a required authorization decorator or a banned crypto helper.
QL is harder to learn and far more expressive. A skilled author can encode invariants that no pattern language can state, and after an incident can describe the exact mistake and run it across every repository. The standard libraries are open, so you can read why a result fired and fix the model instead of suppressing the alert.
This axis favors Semgrep for breadth of authorship and CodeQL for depth of what a single author can express.
Where each one falls short #
Semgrep's month three problem is a ceiling on reasoning. Teams write excellent convention rules, then try to catch a real injection path that crosses files and discover the open source engine will not follow it. At that point you either adopt the commercial engine or accept the gap. The other slow failure is registry sprawl: someone enabled a broad rule pack early, developers learned to ignore the comments, and now the tool has a credibility problem that takes deliberate pruning to fix.
CodeQL's month three problem is that nobody learned QL. The default suites run, produce reasonable results, and the organization never gets the variant hunting or framework modeling that justified choosing it. Meanwhile scan times on large compiled projects stretch pull request feedback, extraction can break when the build changes, and outside GitHub the integration work falls on whoever set it up. Language coverage deserves a check too: PHP appears in Semgrep's language list but not in CodeQL's.
Pick Semgrep if / Pick CodeQL if #
Pick Semgrep if:
- You scan many languages, including ones whose builds are fragile or undocumented.
- Your most common bugs come from misusing internal libraries, wrappers or decorators.
- You want developers, not only the security team, reading and proposing rules.
- You are on GitLab, Jenkins or CircleCI and want a scanner that behaves the same everywhere.
- You need findings on pull requests within minutes of first rollout.
Pick CodeQL if:
- Your repositories are on GitHub and results should land in the security tab developers already see.
- Your highest risk is untrusted input crossing many functions before reaching a sink.
- You have a security engineer prepared to invest in QL and framework models.
- You routinely turn incidents into organization wide searches for the same mistake.
- Your compiled builds are healthy and reproducible in CI.
Frequently asked questions #
Can Semgrep's open source engine replace CodeQL for data flow bugs? Not fully. Its taint mode works within a single file, so flows that cross files or function boundaries need the commercial engine. For convention enforcement and single file taint, the open source engine is often enough.
Do I need to write QL to get value from CodeQL? No, the shipped query suites give a solid baseline scanner. The distinctive value, modeling your frameworks and hunting variants, arrives only when someone writes queries, so plan for engineer time if that is why you chose it.
Is it wasteful to run both? Not if each has a defined job. Keep Semgrep for fast convention checks on every change and CodeQL for scheduled or pull request data flow analysis on critical services, and deduplicate overlapping injection findings so developers do not triage the same issue twice.
Which one is easier to use outside GitHub? Semgrep, because it needs no build and integrates with most CI systems directly. CodeQL runs anywhere through its CLI, but you own the extraction, database management and SARIF handling yourself.
Full profiles: Semgrep and GitHub CodeQL.