Issue:
In Histogram, some values were being dropped when the histogram is accessed concurrently by multiple threads.
Fix:
In getOrCreateHistogram, the code was copying the dictionary into a local variable, and checking it for the label without a lock. If not there, it would take the lock again to add it. But in line 224, after taking the lock again, it was rechecking the local variable, which of course would not have changed in the meantime. But the original dictionary could have changed, and that's the one we should check the second time.
Improved the concurrency test, as it was missing the bug by only checking the total count and sum without labels.
* Introduce PromSummary.capacity to improve performance
* Add a test for summary with custom capacity
* Use Deque instead of Array to store summary values
* Revert "Use Deque instead of Array to store summary values"
This reverts commit 1dd8c32f1a.
* Use CircularBuffer instead of Array to store summary values
* Slightly improve wording in comments
- Updates DimensionLabels' Equatable and Hashable conformances to avoid
allocating unnecessary strings. Updates DimensionHistogramLabels and
DimensionSummaryLabels to use DimensionLabels internally and removes
manually implemented protocol conformances.
Motivation:
Metrics collecting systems should be efficient enough so that their
impact on the system being instrumented is negligible.
Modifications:
- Store `PromMetric`s in a dictionary keyed by label in
`PrometheusClient`; this gives fast lookup when checking if a metric
already exists
- Remove the `metricTypeMap`, we can recover the same information from
`PromMetric`
- Remove calls to `getMetricInstrumet` in `PrometheusMetricsFactory`
these were redundant as the same check is done in `createCounter` etc.
- In each `createCounter` (etc.) call we now hold the lock for the
duration of the call to avoid races between checking the cache and
creating and storing a new metric.
- The `PrometheusLabelSanitizer` now checks whether input needs
santizing instead of santizing all input
- Sanitizing is done in one step, mapping each utf8 code point to a
sanitized code point rather than lowercasing and then checking the
validity of each character.
Result:
- Sanitizing pre-sanitized labels is ~100x faster
- Sanitizing non-sanitized label is ~20x faster
- Incrementing 10 counters is ~20x faster
- Incrementing 100 counters is ~40x faster
- Incrementing 1000 counters is ~250x faster
- (Similar results for other metrics.)
* Add Histogram bucket generation methods
* Make buckets a separate type
* Linux tests
* Make buckets ExpressibleByArrayLiteral
* Add histogram.time()
* Use DispatchTime.now().uptimeNanoseconds instead of Date()
* Apply suggestions from code review
Co-Authored-By: Konrad `ktoso` Malawski <konrad_malawski@apple.com>
* Fix final review remarks
* package.resolved
* WIP: Implement Time Units
* Wrap tests & Up NIO version
* Apply suggestions from code review
Co-Authored-By: Joe Smith <yasumoto7@gmail.com>
* Update linux tests
We might want to instead break this out into two separate functions— one that takes a label, the other that doesn't. That might allow us to then properly pass the callback and only call it once when everything is done.