prometheus

Commit Graph

Author	SHA1	Message	Date
Morten Siebuhr	ffc8cab39a	Updates fuzzers to discard less interesting data	9 years ago
Brian Brazil	ef55fd6176	Add unittest for using a metric for thresholds with group_left.	9 years ago
Morten Siebuhr	981b636004	Bring fuzzer error handling in line.	9 years ago
Morten Siebuhr	9eb2e98509	Fix up documentation + go fmt.	9 years ago
Morten Siebuhr	7371dcc787	Fuzzing corpus for ParseMetric.	9 years ago
Morten Siebuhr	5fec020b27	Initial fuzzing corpus for ParseExpr.	9 years ago
Morten Siebuhr	0ebcca5eb7	Add basic fuzzer of the parser.	9 years ago
Brian Brazil	68e70d992a	Clarify error message around on(x) group_left(x)	9 years ago
Brian Brazil	7201c010c4	Rename On to MatchingLabels	9 years ago
Brian Brazil	d991f0cf47	For many-to-one matches, always copy label from one side. This is a breaking change for everyone using the machine roles labeling approach.	9 years ago
Brian Brazil	768d09fd2a	Change on+group_* to take copy from the one side. If the label doesn't exist on the one side, it's not copied. All labels on the many inside are included, this is a breaking change but likely low impact.	9 years ago
Brian Brazil	d1edfb25b3	Add support for OneToMany with IGNORING. The labels listed in the group_ modifier will be copied from the one side to the many side. It will be valid to specify no labels. This is intended to replace the existing ON/GROUP_* support.,	9 years ago
Brian Brazil	1d08c4fef0	Add 'ignoring' as modifier for binops. Where 'on' uses the given labels to match, 'ignoring' uses all other labels to match. group_left/right is not supported yet.	9 years ago
Brian Brazil	f5084ab1c5	Add tests for group_left/group_right	9 years ago
Fabian Reinartz	fceedfa807	Add error message if old alert rule tokens are read	9 years ago
Julius Volz	6ac39700ea	Fix missing printed keep_common without grouping.	9 years ago
Jonathan Boulle	38098f8c95	Add missing license headers Prometheus is Apache 2 licensed, and most source files have the appropriate copyright license header, but some were missing it without apparent reason. Correct that by adding it.	9 years ago
Tobias Schmidt	8cc86f25c0	Implement relative complement set operator "unless" The `unless` set operator can be used to return all vector elements from the LHS which do not match the elements on the RHS. A use case is to return all metrics for nodes which do not have a specific role: node_load1 unless on(instance) chef_role{role="app"}	9 years ago
Tobias Schmidt	e82ef154ee	Remove unused code leftovers	9 years ago
Tobias Schmidt	4c3dc25e35	Fix whitespace in promql test data	9 years ago
Fabian Reinartz	235e6c554b	Use ContainsRune	9 years ago
eliothedeman	1543ef92b2	Adds holt-winters query function	9 years ago
Fabian Reinartz	ab3d7a0ec0	Remove old alerting syntax	9 years ago
beorn7	4b574e8a61	Switch chunk encoding to type 2 where it was hardcoded type 1 before The chunk encoding was hardcoded there because it mostly doesn't matter what encoding is chosen in that test. Since type 1 is battle-hardened enough, I'm switching to type 2 here so that we can catch unexpected problems as a byproduct. My expectation is that the chunk encoding doesn't matter anyway, as said, but then "unexpected problems" contains the word "unexpected".	9 years ago
Brian Brazil	8788701ce7	Add test for incorrect behaviour	9 years ago
Brian Brazil	39d556f0d5	Move all the operator tests into one file	9 years ago
beorn7	dad302144d	Make a naked return less naked	9 years ago
beorn7	836f1db04c	Improve MetricsForLabelMatchers WIP: This needs more tests. It now gets a from and through value, which it may opportunistically use to optimize the retrieval. With possible future range indices, this could be used in a very efficient way. This change merely applies some easy checks, which should nevertheless solve the use case of heavy rule evaluations on servers with a lot of series churn. Idea is the following: - Only archive series that are at least as old as the headChunkTimeout (which was already extremely unlikely to happen). - Then maintain a high watermark for the last archival, i.e. no archived series has a sample more recent than that watermark. - Any query that doesn't reach to a time before that watermark doesn't have to touch the archive index at all. (A production server at Soundcloud with the aforementioned series churn and heavy rule evaluations spends 50% of its CPU time in archive index lookups. Since rule evaluations usually only touch very recent values, most of those lookup should disappear with this change.) - Federation with a very broad label matcher will profit from this, too. As a byproduct, the un-needed MetricForFingerprint method was removed from the Storage interface.	9 years ago
Patrick Bogen	250344b344	use short variable assignment	9 years ago
Patrick Bogen	2062fbae0f	rewrite operator balancing to be recursive	9 years ago
beorn7	0ea5801e47	Handle errors caused by data corruption more gracefully This requires all the panic calls upon unexpected data to be converted into errors returned. This pollute the function signatures quite lot. Well, this is Go... The ideas behind this are the following: - panic only if it's a programming error. Data corruptions happen, and they are not programming errors. - If we detect a data corruption, we "quarantine" the series, essentially removing it from the database and putting its data into a separate directory for forensics. - Failure during writing to a series file is not considered corruption automatically. It will call setDirty, though, so that a crashrecovery upon the next restart will commence and check for that. - Series quarantining and setDirty calls are logged and counted in metrics, but are hidden from the user of the interfaces in interface.go, whith the notable exception of Append(). The reasoning is that we treat corruption by removing the corrupted series, i.e. a query for it will return no results on its next call anyway, so return no results right now. In the case of Append(), we want to tell the user that no data has been appended, though. Minor side effects: - Now consistently using filepath.* instead of path.*. - Introduced structured logging where I touched it. This makes things less consistent, but a complete change to structured logging would be out of scope for this PR.	9 years ago
beorn7	79a2ae2d2e	Add missing test file	9 years ago
beorn7	2581648f70	Separate iterators by offset Add test that exposes the problem.	9 years ago
Fabian Reinartz	95c9706d2d	Fix missing comment period.	9 years ago
Julius Volz	9ea2465b99	Fix typo in lexer test.	9 years ago
Tobias Schmidt	907b1380a7	Add tests to specify the string escaping behavior	9 years ago
beorn7	c740789ce3	Improve predict_linear Fixes https://github.com/prometheus/prometheus/issues/1401 This remove the last (and in fact bogus) use of BoundaryValues. Thus, a whole lot of unused (and arguably sub-optimal / ugly) code can be removed here, too.	9 years ago
beorn7	454ecf3f52	Rework the way ranges and instants are handled In a way, our instants were also ranges, just with the staleness delta as range length. They are no treated equally, just that in one case, the range length is set as range, in the other the staleness delta. However, there are "real" instants where start and and time of a query is the same. In those cases, we only want to return a single value (the one closest before or at the equal start and end time). If that value is the last sample in the series, odds are we have it already in the series object. In that case, there is no need to pin or load any chunks. A special singleSampleSeriesIterator is created for that. This should greatly speed up instant queries as they happen frequently for rule evaluations.	9 years ago
beorn7	0e202dacb4	Streamline series iterator creation This will fix issue #1035 and will also help to make issue #1264 less bad. The fundamental problem in the current code: In the preload phase, we quite accurately determine which chunks will be used for the query being executed. However, in the subsequent step of creating series iterators, the created iterators are referencing _all_ in-memory chunks in their series, even the un-pinned ones. In iterator creation, we copy a pointer to each in-memory chunk of a series into the iterator. While this creates a certain amount of allocation churn, the worst thing about it is that copying the chunk pointer out of the chunkDesc requires a mutex acquisition. (Remember that the iterator will also reference un-pinned chunks, so we need to acquire the mutex to protect against concurrent eviction.) The worst case happens if a series doesn't even contain any relevant samples for the query time range. We notice that during preloading but then we will still create a series iterator for it. But even for series that do contain relevant samples, the overhead is quite bad for instant queries that retrieve a single sample from each series, but still go through all the effort of series iterator creation. All of that is particularly bad if a series has many in-memory chunks. This commit addresses the problem from two sides: First, it merges preloading and iterator creation into one step, i.e. the preload call returns an iterator for exactly the preloaded chunks. Second, the required mutex acquisition in chunkDesc has been greatly reduced. That was enabled by a side effect of the first step, which is that the iterator is only referencing pinned chunks, so there is no risk of concurrent eviction anymore, and chunks can be accessed without mutex acquisition. To simplify the code changes for the above, the long-planned change of ValueAtTime to ValueAtOrBefore time was performed at the same time. (It should have been done first, but it kind of accidentally happened while I was in the middle of writing the series iterator changes. Sorry for that.) So far, we actively filtered the up to two values that were returned by ValueAtTime, i.e. we invested work to retrieve up to two values, and then we invested more work to throw one of them away. The SeriesIterator.BoundaryValues method can be removed once #1401 is fixed. But I really didn't want to load even more changes into this PR. Benchmarks: The BenchmarkFuzz.* benchmarks run 83% faster (i.e. about six times faster) and allocate 95% fewer bytes. The reason for that is that the benchmark reads one sample after another from the time series and creates a new series iterator for each sample read. To find out how much these improvements matter in practice, I have mirrored a beefy Prometheus server at SoundCloud that suffers from both issues #1035 and #1264. To reach steady state that would be comparable, the server needs to run for 15d. So far, it has run for 1d. The test server currently has only half as many memory time series and 60% of the memory chunks the main server has. The 90th percentile rule evaluation cycle time is ~11s on the main server and only ~3s on the test server. However, these numbers might get much closer over time. In addition to performance improvements, this commit removes about 150 LOC.	9 years ago
Julius Volz	9b6d69610a	Fix various typos in comments. Helpfully reported by https://goreportcard.com/report/github.com/prometheus/prometheus :)	9 years ago
Brian Brazil	9d0112d7cf	Add without aggregator modifier. This has the advantage that the user doesn't need to list all labels they want to keep (as with "by") but without having to worry about inconsistent labels as when there's only one time series (as with "keeping_common"). Almost all aggregation should use this rather than the existing two options as it's much less error prone and easier to maintain due to not having to always add in "job" plus whatever other common job-level labels you have like "region".	9 years ago
Brian Brazil	b7ef0b45e8	Break aggregation tests out. Add missing tests.	9 years ago
beorn7	a7408bfb47	Unify duration parsing It's actually happening in several places (and for flags, we use the standard Go time.Duration...). This at least reduces all our home-grown parsing to one place (in model).	9 years ago
Fabian Reinartz	a6935024e1	Remove old WITH clause in alert printing	9 years ago
Tobias Schmidt	1a91cd6e09	Rename matrix to range selector in external error messages The documentation speaks about range vectors and range vector selectors. This change does not fix all issues, we might still expose the term "Matrix" in error messages using %T.	9 years ago
Tobias Schmidt	411ca4dba1	Consolidate offset modifier parsing Remove duplicated offset modifier parsing and ensure offset can only appear at the end of a selector statement.	9 years ago
Fabian Reinartz	6b4a6962d2	Support old alerting rule syntax	9 years ago
Brian Brazil	c77c3a8c56	promql: Limit extrapolation of delta/rate/increase The new implementation detects the start and end of a series by looking at the average sample interval within the range. If the first (last) sample in the range is more than 1.1*interval distant from the beginning (end) of the range, it is considered the first (last) sample of the series as a whole, and extrapolation is limited to half the interval (rather than all the way to the beginning (end) of the range). In addition, if the extrapolated starting point of a counter (where it is zero) is within the range, it is used as the starting point of the series. Fixes #581	9 years ago
Brian Brazil	89760dd77d	Handle NaN for min/max. Similar to topk and sort, prefer not returning NaN where possible.	9 years ago
Brian Brazil	bac1f28cad	Similar to topk/bottomk, have sort/sort_desc put NaN at end. This makes topk and bottomk consistent with the sorting functions, as per #1271.	9 years ago

... 2 3 4 5 6 ...

341 Commits (d3a1ff1abf9e426c085b8812b467d0f2b48c0a72)