* fix expired Azure secrets being silently dropped
The Azure Entra service principal v2 detector dropped findings entirely
when Azure returned AADSTS7000222 (secret expired). The `continue
SecretLoop` in ProcessData skipped result creation, causing expired
secrets to vanish from output and stop receiving last_seen updates.
Changed the ErrSecretExpired handler to emit an unverified result with
the expiry error preserved, consistent with how ErrSecretInvalid and
ErrConditionalAccessPolicy are handled.
Made-with: Cursor
* Address PR feedback: treat expired secret as definitively invalid
Reviewers noted that an expired secret (AADSTS7000222) is not an
indeterminate verification state — it is definitively invalid. Pass nil
instead of the verification error when creating the result, and update
tests accordingly.
Made-with: Cursor
Return an explicit error when AnalyzePermissions yields nil info
instead of passing nil to secretInfoToAnalyzerResult. Wrap classic
PAT repo/gist enumeration errors with context for easier debugging.
* fix(azure-refresh-token): handle AADSTS50173 as explicit revocation signal
* added same functionality for ErrTokenExpired
* removed TokenLoop since it is no longer used
* removed the createResult to avoid guessing at the secret information.
Triggers on release publish events to run the release bot, which
generates release notes using GitHub, Jira, and AI services.
Adapted from the thog repo workflow with trufflehog-specific adjustments:
repository argument set to trufflehog, environment requirement removed
in favor of a repo-level secret, permissions restricted, and a fork
guard added for consistency with other trufflehog workflows.
Made-with: Cursor
Also adds comments to:
- .goreleaser.yml: explains why make_release is set to false
- .github/workflows/release.yml: document release/artifact state at each step
* use struct-based SourceMetadataFunc signature across git sources
* incorporated feedback
- pass SourceMetadataInfo by value
- remove LegacySourceMetadataFunc
* Fix command injection vulnerability in TUI
The TUI was building a command string from user input via string
concatenation and passing it to `sh -c` through syscall.Exec. This
allowed shell metacharacters in any TUI input field (git URI, file
path, tokens, etc.) to be interpreted as shell commands.
Replace the `sh -c` invocation with a direct syscall.Exec of the
trufflehog binary, passing arguments as a proper argv array. This
eliminates shell interpretation entirely.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Add tilde expansion for TUI args after removing shell layer
Since sh -c was removed to fix command injection, ~/foo paths entered
in the TUI are no longer expanded by a shell. This adds a narrow
expandTilde helper that replaces a leading ~ with os.UserHomeDir()
before exec, restoring path resolution without reintroducing any
shell interpretation.
Guards against empty $HOME to prevent ~/foo silently resolving to /foo.
Made-with: Cursor
---------
Co-authored-by: Bryan Beverly <bryan.beverly@trufflesec.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* Confine symlink state handling to scanSymlink in Filesystem source
* Fix s.canFollowSymlinks snafu
* Update symlink tests to use starting depth 0
* Missed one
* Remove symlink checking from scanFile; this is now always handled in scanSymlink
* Confine errgroup.Groups to scanDir in the Filesystem source (#4808)
* Move path parameter after rootPath parameter in the Filesystem source
* Move the depth parameter too
* Only create an errgroup.Group inside scanDir (where it's used) in the Filesystem source
* Update README formatting and CLI help output
Simplify HTML markup to plain markdown, update the git --help output
to reflect current flags and subcommands, and fix minor formatting
inconsistencies.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Update README: remove $ prompts and refresh CLI help output
Remove leading $ from example commands and replace outdated
trufflehog git --help output with current CLI flags and subcommands.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* [INS-241] New detector (datadogapikey) for datadog apikeys
* Analyzer updated to cater endpoint
* Added new tests for anlyzers
* Removed print statement
* resolved comments and fixed integration tests.
* resolved comments
* changed cli prompt
* Fixed the comments and added app key validation in analyzer
* renamed regex variable
* Added found verified endpoint to ExtraData
* Clean up Analyze function by removing comments
Removed commented-out code for appKey and endpoint.
* [INS-286] Added support to analyze just the apikey in datadog's analyzer
* fixed linter issue
* fixed comment and introduced snake case to make analyzer code cosistent and also fixed flaky tests
* resolved bugbot comment
* updated protos
* resolve conflicts
* updated protobuffs and resolved bugbot comments
* made regex idiomatic
* fixed ssrf vulnerability
* fixed string formatting
* [INS-233] Added support to verify token agains all datadog domains
* [INS-233] Added support to verify token agains all datadog domains
* Fixed cloud endpoint test
* Added precedence to endpoint selection userdefiner -> datafound -> default
* fixed one integration test
* [INS-240] add API key verification fallback when app key verification fails
* Resolved comments
* Fixed the tests according to new changes
* Removed apikey verification logic datadogtoken file
* Reverted the engine test changed earlier
* Resolved comment(s)
* added /api to endpoint
* removed unecessary print
* removed configuredEndpoint to simplify logic
* removed matching with /api suffix
* Fixed the failing integration test
* Update keys in AnalysisInfo map to use snake_case
* fixed bot comments
* fixed ssrf vulnerablity
#4742 (4563dde124) introduced a change to the filesystem source resumption tracking that caused it to start growing linearly with subdirectory count - which causes the payload to get intractably big on large data sets. This commit is an attempt to resolve the issue.
Note that the resumption code still has a bug that can cause data to get inadvertently skipped due to mishandling of the internal parallelization of the scan. This bug has been present for a long time and is present in other sources, so fixing it is out of scope here.
This commit _also_ introduces a new bug related to the fact that lexicographic sorting is not completely appropriate for the resumption check. This needs to be cleaned up as a fast follow, but it's still less serious than the current bug that prevents all scans of large data sets.
I forked go-ldap/ldap from someone else's PR that had partial work to
add contexts, then added just enough to have context on Bind, which
could block if the server does the delay-on-bad-password thing, in
addition to the Dial itself timing out.
there's more room to improve LDAP verification but for now it can't
hold up scanning
* added detector for artifactory reference tokens
* add artifactory reference token detector to the no cloud endpoints list
* address mustansir feedback; remove the invalid host deletion
* use detectors.DetectorHttpClientWithNoLocalAddresses instead of common.SaneHttpClient() just like sibling artifactory detector
* [secret-storage] Thread original chunk data through engine pipeline
Adds OriginalData/ChunkData fields to preserve pre-decode source data
through the scan pipeline:
1. Chunk.OriginalData: captures chunk.Data before iterativeDecode
2. engine.go: sets chunk.OriginalData = chunk.Data before decode
3. ResultWithMetadata.ChunkData: populated by CopyMetadata from
OriginalData (falls back to Data when nil)
This enables downstream consumers (e.g. the dispatcher in thog) to
access the original source data for secret storage encryption.
* Update TestChunkSize for OriginalData field addition
Chunk struct grew from 80 to 104 bytes with the OriginalData []byte
slice header (24 bytes). Field placement is already optimal (adjacent
to Data []byte).
* fix: preserve OriginalData field in EscapedUnicode decoder
The EscapedUnicode decoder constructed a new sources.Chunk manually
copying fields but omitted OriginalData. This caused CopyMetadata to
fall back to the decoded Data instead of the original pre-decode content,
defeating the purpose of preserving original chunk data for secret
storage encryption.
* Address PR review feedback: use testify/assert, add nil-guard comment, remove stale alignment comment
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Chunk.Verify is an odd field - it originally conveys whether a source is going to run with verification, but then, at a certain point in the scanning pipeline, is mutated such that it instead indicates whether the chunk should be scanned with verification - which is not solely dependent on the source's verify flag. This is unnecessarily difficult to understand and maintain. This commit separates those two pieces of information into two flags:
- Chunk.Verify has been renamed to Chunk.SourceVerify
- It is no longer mutated; instead "should this chunk's secrets be verified?" is now captured by a new field on detectableChunk
* enabled symlinks with maximum depth support
* resolved concurrency bugs and added maxDepthOption to cli
* Removed visited path map and tracked symlink depth by maintaining a counter variable
* separated symlink scanning from scanDir
* resolved bugbot comments
* introduced hash as a separator to avoid collisions
* added rotation on 403s, it appears that this detector considered them to be indeterminate failures
* tightened logic for rotation
* Apply suggestion from @nabeelalam
Co-authored-by: Nabeel Alam <nabeelalam811@gmail.com>
* Modify error return for RabbitMQ access refusal
Update error handling to return nil instead of the error for specific access issues.
---------
Co-authored-by: Kashif Khan <70996046+kashifkhan0771@users.noreply.github.com>
Co-authored-by: Nabeel Alam <nabeelalam811@gmail.com>
* fix(ftp): set read deadline on connection to prevent indefinite hang
The FTP detector uses ftp.DialWithTimeout to bound the TCP connection,
but after the connection is established the banner read and login have
no deadline. If a server accepts TCP but never sends a response, the
detector blocks indefinitely, stalling the entire scan pipeline.
Switch to DialWithDialFunc and set a deadline on the connection that
covers the full handshake (banner + login).
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ftp): check error return from conn.SetDeadline
Co-authored-by: Dylan Ayrey <dxa4481@rit.edu>
---------
Co-authored-by: Dylan Ayrey <dylan@Dylans-MacBook-Pro.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Dylan Ayrey <dxa4481@rit.edu>
* Replace logConfig with zapcore.Core
* Add SyncFunc type and tests
* Change Sentry functions to accept a *sentry.Client
This allows the caller to do the error handling.
* Rename newCoreConfig to newCore
* Remove redundant WithCore option
* Update doc comments
* add google gemini api key detector
* change detector name to google cloud api key, mark as verified if 403 is returned
* add build tags for integration test
* changes in defaults.go
* Revert "changes in defaults.go"
This reverts commit 12e7b6f4aa.
* Revert "change detector name to google cloud api key, mark as verified if 403 is returned"
This reverts commit e46bb29b40.
* revert google cloud api changes, change keyword to gemini, add extra field active_google_key
* use aizasy as keyword instead of gemini
* close response body after draining
* remove \b from regex to support keys that end with -
* add \b to the beginning
* feat: iterative decoding pipeline with configurable depth
Decoders (base64, UTF-16, escaped unicode) now chain iteratively:
each decoder's output is fed back through all decoders until no new
transformations occur or --max-decode-depth is reached (default: 5).
This finds secrets hidden inside layered encodings, e.g. a base64
Docker auth blob containing a GCP private key, or a UTF-16 file
with base64-encoded credentials.
At depth=1 behavior is identical to the previous implementation.
Extra depths exit early when no new data is produced, so the cost
is <5% wall time on a large repo scan.
Co-authored-by: Dylan Ayrey <dxa4481@rit.edu>
* docs: iterative decoding performance data
Co-authored-by: Dylan Ayrey <dxa4481@rit.edu>
* comment: explain why PLAIN decoder is skipped at depth > 0
Co-authored-by: Dylan Ayrey <dxa4481@rit.edu>
* refactor: extract iterativeDecode, address review feedback
- Extract decode loop into standalone iterativeDecode() function,
separating decoding from channel dispatch (rosecodym, camgunz).
- Drop decodeInput struct, use []byte directly (camgunz).
- Remove redundant maxDepth clamp from scannerWorker (camgunz).
- Inline decoderType variable (camgunz).
- Replace byteSliceSeen with slices.ContainsFunc (camgunz).
Co-authored-by: Dylan Ayrey <dxa4481@rit.edu>
* fix: remove unused decodeLatency metric (lint)
Co-authored-by: Dylan Ayrey <dxa4481@rit.edu>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Our underlying "ubuntu-latest" image moved to Docker v29, which defaults
to a containerd image store, which can store provenance info and creates
even single images with a manifest by default, but goreleaser with our
current config isn't expecting that. This unblocks the release build,
but I believe moving to dockers_v2 will fully resolve.
See also: docker/build-push-action#755