What we still need to verify : 2 points in this profile are not yet confirmed against vendor documentation.
- Language support tiers vary by maturity level per language, confirm the current matrix
- Which capabilities sit in the open source engine versus the commercial engine changes over time, confirm
Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.
What it does
Semgrep matches rules against a parsed representation of your code. A rule is written as a fragment of the target language with metavariables standing in for the parts that vary, so a rule to find a shell execution with a non literal argument looks like the call it is meant to catch. Matching happens on the syntax tree, so whitespace, comments and argument formatting do not defeat a rule the way they defeat a regular expression. Rules compose through operators for containment, exclusion and metavariable comparison, and are expressed in YAML alongside metadata for severity, CWE mapping and a fix suggestion.
Beyond single pattern matching there is taint mode, where a rule declares sources, sinks, sanitizers and propagators, and the engine reports untrusted data reaching a dangerous operation. In the open source engine this analysis works within a file. The commercial engine extends it across files and across function boundaries and adds reachability judgement for dependency findings. Rules come from a public registry organized by language, framework and vulnerability class, and you can mix registry rules with your own in the same run.
Where it fits
It runs anywhere source is available: a developer laptop, a pre-commit hook, a pull request job, or a scheduled full repository scan. No build and no dependency installation are required, which is the main reason it spreads quickly across a polyglot estate. The security team typically owns rule curation and writes organization specific rules, for example forbidding a deprecated internal crypto wrapper or requiring a particular authorization decorator, while developers see the results as pull request comments. Diff aware scanning, which reports only findings introduced by the change, is what makes it usable on a large legacy codebase.
Strengths
- Rule syntax that reads like the code it matches, so developers can read, review and write rules without learning a query language.
- No build step, so onboarding a new repository takes minutes regardless of language.
- Diff aware scanning makes adoption on an existing codebase practical.
Limitations
- The open source engine analyzes within a file. Cross file taint tracking sits in the commercial product, and that gap covers real vulnerability classes.
- Registry rule quality is uneven. Community rules vary in precision and you should curate rather than enable everything.
- Syntactic matching without semantic context produces false positives where safety depends on something the rule cannot see, such as a value validated three frames up the stack.
Who it suits
A strong default for teams with many languages and a security engineer willing to write rules. Less compelling if you want deep interprocedural analysis without the commercial engine, or if you need a turnkey scanner with no rule curation at all.
Used Semgrep? Recommend it under your own name and title.
Recommend this tool