prometheus

Commit Graph

Author	SHA1	Message	Date
AllenZMC	5c2c9a03e9	fix word 'consequentally' to 'consequently' (#5827 ) Signed-off-by: czm <zhongming.chang@daocloud.io>	5 years ago
beorn7	dd81912554	Add objectives to Summaries With the next release of client_golang, Summaries will not have objectives by default. To not lose the objectives we have right now, explicitly state the current default objectives. Signed-off-by: beorn7 <beorn@grafana.com>	6 years ago
pbhudiaBAE	43953b105b	Sorting alerts by group name in /alerts (#5448 ) * Working group name Signed-off-by: Pritam Bhudia <pritam.bhudia@baesystems.com> * Working categorised by group name Signed-off-by: Pritam Bhudia <pritam.bhudia@baesystems.com> * Changed group sorting in web Signed-off-by: Pritam Bhudia <pritam.bhudia@baesystems.com> * Fixed group sorting and comments Signed-off-by: Pritam Bhudia <pritam.bhudia@baesystems.com> * Fixed group sorting and comments with gofmt Signed-off-by: Pritam Bhudia <pritam.bhudia@baesystems.com> * Added file and group name Signed-off-by: Pritam Bhudia <pritam.bhudia@baesystems.com> * reverted back to full path to yml file Signed-off-by: Pritam Bhudia <pritam.bhudia@baesystems.com>	6 years ago
Yao Zengzeng	5544cb252a	fix some mistakes in comments (#5533 ) Signed-off-by: YaoZengzeng <yaozengzeng@zju.edu.cn>	6 years ago
Simon Pasquier	45506841e6	*: enable all default linters (#5504 ) Signed-off-by: Simon Pasquier <spasquie@redhat.com>	6 years ago
Bjoern Rabenstein	76102d570c	Add test for external labels in label template Signed-off-by: Bjoern Rabenstein <bjoern@rabenste.in>	6 years ago
Bjoern Rabenstein	38d518c0fe	Rework #5009 after comments Signed-off-by: Bjoern Rabenstein <bjoern@rabenste.in>	6 years ago
Sylvain Rabot	335a34486e	Add external labels to template expansion This affects the expansion of templates in alert labels and annotations and console templates. Signed-off-by: Sylvain Rabot <sylvain@abstraction.fr>	6 years ago
Tariq Ibrahim	8fdfa8abea	refine error handling in prometheus (#5388 ) i) Uses the more idiomatic Wrap and Wrapf methods for creating nested errors. ii) Fixes some incorrect usages of fmt.Errorf where the error messages don't have any formatting directives. iii) Does away with the use of fmt package for errors in favour of pkg/errors Signed-off-by: tariqibrahim <tariq181290@gmail.com>	6 years ago
James Ravn	e15d8c5802	reload: copy state on both name and labels (#5368 ) * reload: copy state on both name and labels Fix https://github.com/prometheus/prometheus/issues/5193 Using just name causes the linked issue - if new rules are inserted with the same name (but different labels), the reordering will cause stale markers to be inserted in the next eval for all shifted rules, despite them not being stale. Ideally we want to avoid stale markers for time series that still exist in the new rules, with name and labels being the unique identifer. This change adds labels to the internal map when copying the old rule data to the new rule data. This prevents the problem of staling rules that simply shifted order. If labels change, it is a new time series and the old series will stale regardless. So it should be safe to always match on name and labels when copying state. Signed-off-by: James Ravn <james@r-vn.org>	6 years ago
David Symonds	46361a7c85	rules: Fix sorting of result from (*Manager).RuleGroups (#5260 ) The previous code was defective in that it never sorted groups within a file due to doing a multi-key sort incorrectly. Signed-off-by: David Symonds <dsymonds@gmail.com>	6 years ago
beorn7	2db1eeb4ec	Fix prometheus_rule_group_last_evaluation_timestamp_seconds It should be a unix timestamp, not the seconds in the minute. Signed-off-by: beorn7 <beorn@soundcloud.com>	6 years ago
Ganesh Vernekar	787eb1e904	Set rule_group_last_duration_seconds to seconds (#5153 ) Signed-off-by: Ganesh Vernekar <cs15btech11018@iith.ac.in>	6 years ago
Matt Layher	302148fd69	*: apply gofmt -s Signed-off-by: Matt Layher <mdlayher@gmail.com>	6 years ago
Vishnunarayan K I	fd3ef6ba34	Add metric rule_group_rules_loaded to get the number of rules loaded (#5090 ) Signed-off-by: Vishnunarayan K I <appukuttancr@gmail.com>	6 years ago
Simon Pasquier	f678e27eb6	: use latest release of staticcheck (#5057 ) : use latest release of staticcheck It also fixes a couple of things in the code flagged by the additional checks. Signed-off-by: Simon Pasquier <spasquie@redhat.com> Use official release of staticcheck Also run 'go list' before staticcheck to avoid failures when downloading packages. Signed-off-by: Simon Pasquier <spasquie@redhat.com>	6 years ago
Tom Wilkie	121603c417	Expose rules.NewGroupMetrics and rules.Metrics. (#5059 ) Signed-off-by: Tom Wilkie <tom.wilkie@gmail.com>	6 years ago
Tom Wilkie	6e08029b56	Move err to be the last return value from storage.Select. (#5054 ) Signed-off-by: Tom Wilkie <tom.wilkie@gmail.com>	6 years ago
Bartek Płotka	de213d4a5e	rule manager: Moved metric registration to custom registerer which is already available. (#4961 ) Signed-off-by: Bartek Plotka <bwplotka@gmail.com>	6 years ago
AixesHunter	fb8479a677	Variable 'labels' collides with imported package name (#5012 ) Signed-off-by: aixeshunter <aixeshunter@gmail.com>	6 years ago
mknapphrt	f0e9196dca	Return warnings on a remote read fail (#4832 ) Signed-off-by: Mark Knapp <mknapp@hudson-trading.com>	6 years ago
Krasi Georgiev	0754e5334b	querier for RestoreForState not closed. (#4922 ) Signed-off-by: Krasi Georgiev <kgeorgie@redhat.com>	6 years ago
Ben Kochie	c6399296dc	Fix spelling/typos (#4921 ) * Fix spelling/typos Fix spelling/typos reported by codespell/misspell. * UK -> US spelling changes. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Wei Guo	e329cbf673	Add metric prometheus_rule_group_last_evaluation for recording and alerting (#4852 ) * add metric prometheus_rule_group_last_evaluation for recording and alerting Signed-off-by: Wei Guo <me@imkira.com> * fix issues from comments Signed-off-by: Wei Guo <me@imkira.com>	6 years ago
Will Hegedus	193ebe7e34	Updates to /targets and /rules (scrape duration, last evaluation time) (#4722 ) * Add evaluationTimestamp (Last Evaluation) column to display on /rules Signed-off-by: Will Hegedus <wbhegedus@liberty.edu> * Add lastScrapeDuration ("Scrape Duration") to display on /targets Signed-off-by: Will Hegedus <wbhegedus@liberty.edu> * Updates based on Julius' feedback Signed-off-by: Will Hegedus <wbhegedus@liberty.edu> * Update to set timestamp to when eval started (after eval completes) Signed-off-by: Will Hegedus <wbhegedus@liberty.edu> * Update /rules to display time since last evaluation Signed-off-by: Will Hegedus <wbhegedus@liberty.edu> * Re-order Last Eval/Eval Time to be consistent with targets page Signed-off-by: Will Hegedus <wbhegedus@liberty.edu>	6 years ago
Callum Styan	9bca041285	WIP: keep track of samples per query, set a max # of samples (#4513 ) * keep track of samples per query, set a max # of samples that can be in memory at once Signed-off-by: Callum Styan <callumstyan@gmail.com>	6 years ago
Ganesh Vernekar	5790d23fd8	Unit testing for rules (#4350 ) * Unit testing for rules * Specifying order of group evaluation in unit tests Signed-off-by: Ganesh Vernekar <cs15btech11018@iith.ac.in>	6 years ago
Ganesh Vernekar	05726c5ea2	Test template expansion while loading groups (#4537 ) Signed-off-by: Ganesh Vernekar <cs15btech11018@iith.ac.in>	6 years ago
Chris Marchbanks	63ed9d1b70	Send EndsAt along with alerts (#4550 ) Signed-off-by: Chris Marchbanks <csmarchbanks@gmail.com>	6 years ago
Chris Marchbanks	87f1dad16d	throttle resends of alerts to 1 minute by default (#4538 ) Signed-off-by: Chris Marchbanks <csmarchbanks@gmail.com>	6 years ago
Goutham Veeramachaneni	f3b7c22827	rules: add comment about lock taking (#4525 ) Signed-off-by: Goutham Veeramachaneni <gouthamve@gmail.com>	6 years ago
Ganesh Vernekar	c663477688	Fixed TestUpdate in rules/manager_test.go (#4516 ) Signed-off-by: Ganesh Vernekar <cs15btech11018@iith.ac.in>	6 years ago
Julius Volz	8fbe1b5133	Handle a bunch of unchecked errors (#4461 ) There are many more (mostly finalizers like Close/Stop/etc.), but most of the others seemed like one couldn't do much about them anyway. Signed-off-by: Julius Volz <julius.volz@gmail.com>	6 years ago
Ganesh Vernekar	a0a9e7df91	Fix TestForStateRestore (#4476 ) (#4512 ) Signed-off-by: Ganesh Vernekar <cs15btech11018@iith.ac.in>	6 years ago
Julien Pivotto	0b4d22b245	rules/manager: remove a no-longer-relevant comment (#4503 ) Signed-off-by: Julien Pivotto <roidelapluie@inuits.eu>	6 years ago
Chris Marchbanks	11155c7028	Existing alert labels will update based on templates (#4500 ) Signed-off-by: Chris Marchbanks <csmarchbanks@gmail.com>	6 years ago
Fabian Reinartz	b7e2f407de	rules: Fix double-locking of mutex Signed-off-by: Fabian Reinartz <freinartz@google.com>	6 years ago
Benji Visser	8bb6e0dd6e	Show rule evaluation errors on rules page (#4457 ) * adding information about the health and errors for Rules adding Health() and LastError() to the Rule interface. This will allow us to easily surface information about rules. Signed-off-by: noqcks <benny@noqcks.io> * updating rules.html with fields for Rule errors and health state Signed-off-by: noqcks <benny@noqcks.io> * fix code comment grammar & access Rule health/error info using a mutex Signed-off-by: noqcks <benny@noqcks.io> * s/Errors/Error/ in rules.html to remain consistent with targets.html Signed-off-by: noqcks <benny@noqcks.io> * adding periods to code comments in reporting/alerting Signed-off-by: noqcks <benny@noqcks.io> * putting health/error below mutex in struct field Signed-off-by: noqcks <benny@noqcks.io>	6 years ago
Julius Volz	2b8fc062a8	rules: HTML-escape rule YAML marshal errors (#4464 ) This was pointed out by `gosec`. Signed-off-by: Julius Volz <julius.volz@gmail.com>	6 years ago
Julius Volz	90521a65f8	Remove error return value from NotifyFunc() (#4459 ) It's always nil and we also forgot to check it. Signed-off-by: Julius Volz <julius.volz@gmail.com>	6 years ago
Ganesh Vernekar	f1db699dff	Persist alert 'for' state across restarts (#4061 ) Signed-off-by: Ganesh Vernekar <cs15btech11018@iith.ac.in>	6 years ago
Max Leonard Inden	71fafad099	api/v1: Coninue work exposing rules and alerts Signed-off-by: Max Leonard Inden <IndenML@gmail.com>	6 years ago
mg03	31f8ca0dfb	api v1 alerts/rules json endpoint Signed-off-by: mg03 <mgeng03@gmail.com>	6 years ago
Bryan Boreham	afdb66dfac	Expose Group.CopyState() (#4304 ) This makes the `rules` package more useful to projects that use Prometheus as a library. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	6 years ago
Julius Volz	9e3171f6e3	rules: Minor naming/comment cleanups (#4328 ) Signed-off-by: Julius Volz <julius.volz@gmail.com>	6 years ago
Bryan Boreham	2bd510a63e	Make TestUpdate() do some work (#4306 ) Previously it would set no preconditions and check no postconditions, as the `groups` member was empty. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	7 years ago
Alin Sinpalean	9dc763cc03	Run rule evaluation with timestamps precisely evaluation_interval apart (#4201 ) * Run rule evaluation with timestamps precisely evaluation_interval apart from one another. Signed-off-by: Alin Sinpalean <alin.sinpalean@gmail.com>	7 years ago
Mario Trangoni	464e747f1e	fix some comments typos (#4059 )	7 years ago
Bryan Boreham	93494d8b7e	Add an OpenTracing span for each rule (#4027 ) * Add an OpenTracing span for each rule So that tags and child spans can be traced back to the rule that they refer to.	7 years ago
ferhat elmas	ec8e4d8a7c	all: remove unnecessary type conversions (#3992 ) excep promql due to not to create conflict with #3966.	7 years ago
Warren Fernandes	58e2a31db8	Cleans up test by removing unused function (#3969 )	7 years ago
ferhat elmas	ffa673f7d8	General simplifications (#3887 ) Another try as in #1516	7 years ago
Fabian Reinartz	7ccd4b39b8	*: implement query params This adds a parameter to the storage selection interface which allows query engine(s) to pass information about the operations surrounding a data selection. This can for example be used by remote storage backends to infer the correct downsampling aggregates that need to be provided.	7 years ago
Simon Pasquier	81c0ab69e0	Don't reset FiredAt for inactive alerts Otherwise AlertManager receives resolved alerts where StartsAt is zero which fails the validation.	7 years ago
Brian Brazil	30b4439bbd	Remove rule_type label from rule metrics. This is not really needed now that we have rule groups to distinguish rules.	7 years ago
Brian Brazil	b97f4cf48c	Add metrics for rule group interval and last duration.	7 years ago
Brian Brazil	0a42a9fc8f	Copy over rule group duration on reload. This is currently getting lost, this will soon be in a metric and we don't want it dropping to 0 on every reload.	7 years ago
Brian Brazil	aa370fa568	Clarify metric names around rule groups. Make it clear they're about overall rule groups.	7 years ago
Fabian Reinartz	62461379b7	rules: decouple notifier packages The dependency on the notifier packages caused a transitive dependency on discovery and with that all client libraries our service discovery uses.	7 years ago
Fabian Reinartz	4d964a0a0d	rules: make glob expansion a concern of main	7 years ago
Fabian Reinartz	bd9f7460eb	rules: remove config package dependency	7 years ago
Fabian Reinartz	2d0e3746ac	rules: remove dependency on promql.Engine	7 years ago
Fabian Reinartz	2ec5965b75	Merge pull request #3508 from prometheus/uptsdb update TSDB	7 years ago
Fabian Reinartz	83cd270ea4	*: adapt to storage interface changes	7 years ago
Goutham Veeramachaneni	a880c86375	Fix unexported method on exported interface. Also move to model.Duration Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	7 years ago
conorbroderick	55aaece116	Add rule evaluation time	7 years ago
Goutham Veeramachaneni	e1117715fe	rules: remove skipped iterations cuz no throttling Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	7 years ago
Jorge Hernández	6cd0f63eb1	Use testutil in rules subpackage (#3278 ) * Use testutil in rules subpackage * Fix manager test * Use testutil in rules subpackage * Fix manager test * Fix rebase * Change to testutil for applyConfig tests	7 years ago
Krasi Georgiev	e86d82ad2d	Fix regression of alert rules state loss on config reload. (#3382 ) * incorrect map name for the group prevented copying state from existing alert rules on config reload * applyConfig test * few nits * nits 2	7 years ago
Julius Volz	099df0c5f0	Migrate "golang.org/x/net/context" -> "context" (#3333 ) In some places, where ctxhttp or gRPC are concerned, we still need to use the old contexts.	7 years ago
Brian Brazil	cc5499fcad	Only close after checking for err.	7 years ago
Brian Brazil	ee88f0d222	Ensure all values are used or _	7 years ago
Fabian Reinartz	2d0b8e8b94	Merge branch 'master' into dev-2.0	7 years ago
Julius Volz	f7e8348a88	Re-add contexts to storage.Storage.Querier() (#3230 ) * Re-add contexts to storage.Storage.Querier() These are needed when replacing the storage by a multi-tenant implementation where the tenant is stored in the context. The 1.x query interfaces already had contexts, but they got lost in 2.x. * Convert promql.Engine to use native contexts	7 years ago
beorn7	c2e9a151ab	Make all rule links link to the "Console" tab rather than "Graph" Clicking on a rule, either the name or the expression, opens the rule result (or the corresponding expression, repsectively) in the expression browser. This should by default happen in the console tab, as, more often than not, displaying it in the graph tab runs into a timeout.	7 years ago
Fabian Reinartz	d21f149745	*: migrate to go-kit/log	7 years ago
Goutham Veeramachaneni	e1fc9dc78d	Move /rules to new format (#2901 ) Fixes #2891 Signed-off-by: Goutham Veeramachaneni <goutham@boomerangcommerce.com>	7 years ago
Goutham Veeramachaneni	37e7b69f56	Merge remote-tracking branch 'upstream/dev-2.0' into rulegroups Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	c472316fb3	Check done before every rule evaluation. Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	6b70a4d850	Incorporate PR feedback * Move fingerprint to Hash() * Move away from tsdb.MultiError * 0777 -> 0666 for files * checkOverflow of extra fields Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	507790a357	Rework logging to use explicitly passed logger Mostly cleaned up the global logger use. Still some uses in discovery package. Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	dc69645e92	Move back to go-yaml Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	5ff283a7b7	Reflect the grouping in the UI Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	8cca666cf2	Add file name to group. Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	e893c89333	Validate labels and annotations Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	a48a018368	Make sure groups are unique in a single file Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	cea1e99f78	Add update-rules command to promtool Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	e8f55669ea	Move rules to new format Signed-off-by: Goutham Veeramachaneni <goutham@boomerangcommerce.com>	8 years ago
Brian Brazil	dcea3e4773	Don't append a 0 when alert is no longer pending/firing With staleness we no longer need this behaviour.	8 years ago
Brian Brazil	cc867dae60	Copy previous series and alert state more intelligently. Usually rules don't more around, and if they do it's likely that rules/alerts with the same name stay in the same order. If rules/alerts with the same name are added/removed this could cause a blip for one cycle, but this is unavoidable without requiring rule and alert names to be unique - which we don't want to do.	8 years ago
Brian Brazil	9bc68db7e6	Track staleness per rule rather than per group.	8 years ago
Brian Brazil	0451d6d31b	Add unittest for rule staleness, and rules generally.	8 years ago
Brian Brazil	0400f3cfd2	Very basic staleness handling for rules.	8 years ago
Fabian Reinartz	06c2b76cd4	Merge branch 'master' into uptsdb	8 years ago
Alexey Palazhchenko	b0e1ea7c6c	Simplify code, fix typos. (#2719 )	8 years ago
Julius Volz	ac203ef0ee	Add externalURL template function (#2716 ) This allows users to e.g. add links back to the generating Prometheus right in their alert templates.	8 years ago
Julius Volz	fe11c5933a	Fix mutation of active alert elements by notifier (#2656 ) This caused the external label application in the notifier to bleed back into the rule manager's active alerting elements.	8 years ago
Fabian Reinartz	8ffc851147	Merge branch 'master' into dev-2.0	8 years ago
Tobias Schmidt	eaf33759fb	Register forgotten prometheus_evaluator_iterations_total metric	8 years ago
Tobias Schmidt	aaaba57184	Export number of missed rule evaluations In case the execution of all rules takes longer than the configured rule evaluation interval, one or more iterations will be skipped. This needs to be visible to the opterator.	8 years ago
Fabian Reinartz	5772f1a7ba	retrieval/storage: adapt to new interface This simplifies the interface to two add methods for appends with labels or faster reference numbers.	8 years ago
Fabian Reinartz	ad9bc62e4c	storage: extend appender and adapt it	8 years ago
Fabian Reinartz	e94b0899ee	rules: fix tests, remove model types	8 years ago
Fabian Reinartz	f8fc1f5bb2	*: migrate ingestion to new batch Appender	8 years ago
Fabian Reinartz	fecf9532b9	*: fix misc compile errors	8 years ago
Fabian Reinartz	622ece6273	*: fix recording tests, migrate matcher types	8 years ago
Fabian Reinartz	5817cb5bde	: migrate from model. to promql.* types	8 years ago
Fabian Reinartz	e68a3cf21f	rules: update annotations on each iteration	8 years ago
Jonathan Lange	d78dd3593d	Set evaluation interval on Group construction Prevents having object in invalid state, and allows users of public API to construct valid Groups.	8 years ago
Jonathan Lange	31fc357cd8	Make NewGroup and Group.Eval public Allows callers to execute evaluate lists of rules without first writing them to disk.	8 years ago
Jonathan Lange	2a2da40223	Make rule evaluation publicly available Means that a third-party can parse rules and run them with their own execution model.	8 years ago
Matt Bostock	926a5ab3dd	rules/manager.go: Fix race between reload and stop On one relatively large Prometheus instance (1.7M series), I noticed that upgrades were frequently resulting in Prometheus undergoing crash recovery on start-up. On closer examination, I found that Prometheus was panicking on shutdown. It seems that our configuration management (or misconfiguration thereof) is reloading Prometheus then immediately restarting it, which I suspect is causing this race: Sep 21 15:12:42 host systemd[1]: Reloading prometheus monitoring system. Sep 21 15:12:42 host prometheus[18734]: time="2016-09-21T15:12:42Z" level=info msg="Loading configuration file /etc/prometheus/config.yaml" source="main.go:221" Sep 21 15:12:42 host systemd[1]: Reloaded prometheus monitoring system. Sep 21 15:12:44 host systemd[1]: Stopping prometheus monitoring system... Sep 21 15:12:44 host prometheus[18734]: time="2016-09-21T15:12:44Z" level=warning msg="Received SIGTERM, exiting gracefully..." source="main.go:203" Sep 21 15:12:44 host prometheus[18734]: time="2016-09-21T15:12:44Z" level=info msg="See you next time!" source="main.go:210" Sep 21 15:12:44 host prometheus[18734]: time="2016-09-21T15:12:44Z" level=info msg="Stopping target manager..." source="targetmanager.go:90" Sep 21 15:12:52 host prometheus[18734]: time="2016-09-21T15:12:52Z" level=info msg="Checkpointing in-memory metrics and chunks..." source="persistence.go:548" Sep 21 15:12:56 host prometheus[18734]: time="2016-09-21T15:12:56Z" level=warning msg="Error on ingesting out-of-order samples" numDropped=1 source="scrape.go:467" Sep 21 15:12:56 host prometheus[18734]: time="2016-09-21T15:12:56Z" level=error msg="Error adding file watch for \"/etc/prometheus/targets\": no such file or directory" source="file.go:84" Sep 21 15:12:56 host prometheus[18734]: time="2016-09-21T15:12:56Z" level=error msg="Error adding file watch for \"/etc/prometheus/targets\": no such file or directory" source="file.go:84" Sep 21 15:13:01 host prometheus[18734]: time="2016-09-21T15:13:01Z" level=info msg="Stopping rule manager..." source="manager.go:366" Sep 21 15:13:01 host prometheus[18734]: time="2016-09-21T15:13:01Z" level=info msg="Rule manager stopped." source="manager.go:372" Sep 21 15:13:01 host prometheus[18734]: time="2016-09-21T15:13:01Z" level=info msg="Stopping notification handler..." source="notifier.go:325" Sep 21 15:13:01 host prometheus[18734]: time="2016-09-21T15:13:01Z" level=info msg="Stopping local storage..." source="storage.go:381" Sep 21 15:13:01 host prometheus[18734]: time="2016-09-21T15:13:01Z" level=info msg="Stopping maintenance loop..." source="storage.go:383" Sep 21 15:13:01 host prometheus[18734]: panic: close of closed channel Sep 21 15:13:01 host prometheus[18734]: goroutine 7686074 [running]: Sep 21 15:13:01 host prometheus[18734]: panic(0xba57a0, 0xc60c42b500) Sep 21 15:13:01 host prometheus[18734]: /usr/local/go/src/runtime/panic.go:500 +0x1a1 Sep 21 15:13:01 host prometheus[18734]: github.com/prometheus/prometheus/rules.(Manager).ApplyConfig.func1(0xc6645a9901, 0xc420271ef0, 0xc420338ed0, 0xc60c42b4f0, 0xc6645a9900) Sep 21 15:13:01 host prometheus[18734]: /home/build/packages/prometheus/tmp/build/gopath/src/github.com/prometheus/prometheus/rules/manager.go:412 +0x3c Sep 21 15:13:01 host prometheus[18734]: created by github.com/prometheus/prometheus/rules.(Manager).ApplyConfig Sep 21 15:13:01 host prometheus[18734]: /home/build/packages/prometheus/tmp/build/gopath/src/github.com/prometheus/prometheus/rules/manager.go:423 +0x56b Sep 21 15:13:03 host systemd[1]: prometheus.service: main process exited, code=exited, status=2/INVALIDARGUMENT	8 years ago
Julius Volz	c187308366	storage: Contextify storage interfaces. This is based on https://github.com/prometheus/prometheus/pull/1997. This adds contexts to the relevant Storage methods and already passes PromQL's new per-query context into the storage's query methods. The immediate motivation supporting multi-tenancy in Frankenstein, but this could also be used by Prometheus's normal local storage to support cancellations and timeouts at some point.	8 years ago
Julius Volz	ed5a0f0abe	promql: Allow per-query contexts. For Weaveworks' Frankenstein, we need to support multitenancy. In Frankenstein, we initially solved this without modifying the promql package at all: we constructed a new promql.Engine for every query and injected a storage implementation into that engine which would be primed to only collect data for a given user. This is problematic to upstream, however. Prometheus assumes that there is only one engine: the query concurrency gate is part of the engine, and the engine contains one central cancellable context to shut down all queries. Also, creating a new engine for every query seems like overkill. Thus, we want to be able to pass per-query contexts into a single engine. This change gets rid of the promql.Engine's built-in base context and allows passing in a per-query context instead. Central cancellation of all queries is still possible by deriving all passed-in contexts from one central one, but this is now the responsibility of the caller. The central query context is now created in main() and passed into the relevant components (web handler / API, rule manager). In a next step, the per-query context would have to be passed to the storage implementation, so that the storage can implement multi-tenancy or other features based on the contextual information.	8 years ago
beorn7	75bae065fd	Revert "Modify tests to adjust to reverting the /graph changes" This reverts commit `f1ea5bf232`. Part two necessary for reverting the /graph revert.	8 years ago
beorn7	f1ea5bf232	Modify tests to adjust to reverting the /graph changes These tests have been added after the /graph changes and therefore already test the new syntax. This commit has to be reverted together with the previous one to get back to the old new state. sigh	8 years ago
Julius Volz	fe7b8b7fd1	Add missing license header to alerting_test.go	8 years ago
Julius Volz	da7206ec29	Fix rule HTML escaping issues This was mentioned as part of https://github.com/prometheus/alertmanager/issues/452	8 years ago
Brian Brazil	6fc88d4b4d	Remove __name__ from alerts sent to AM. Fixes #1861	8 years ago
Dmitry Vorobev	273e457da4	web: return status code and error message for config resource	8 years ago
Brian Brazil	0509b0f2db	Expand alert templates at eval time. Fixes #1678 #1677	8 years ago
beorn7	064b57858e	Consistently use the `Seconds()` method for conversion of durations This also fixes one remaining case of recording integral numbers of seconds only for a metric, i.e. this will probably fix #1796.	9 years ago
Fabian Reinartz	f7ed2ff706	Merge pull request #1644 from prometheus/beorn7/logging Add missing logging of out-of-order samples	9 years ago
beorn7	b95c096a45	Fix style issues in rules/...	9 years ago
beorn7	45e5775f9b	Add missing logging of out-of-order samples So far, out-of-order samples during rule evaluation were not logged, and neither scrape health samples. The latter are unlikely to cause any errors. That's why I'm logging them always now. (It's alway highly irregular should it happen.) For rules, I have used the same plumbing as for samples, just with a different wording in the message to mark them as a result of rule evaluation.	9 years ago
beorn7	4b574e8a61	Switch chunk encoding to type 2 where it was hardcoded type 1 before The chunk encoding was hardcoded there because it mostly doesn't matter what encoding is chosen in that test. Since type 1 is battle-hardened enough, I'm switching to type 2 here so that we can catch unexpected problems as a byproduct. My expectation is that the chunk encoding doesn't matter anyway, as said, but then "unexpected problems" contains the word "unexpected".	9 years ago
Fabian Reinartz	d89c254849	Make copying alerting state safer. This considers static labels in the equality of alerts to avoid falsely copying state from a different alert definition with the same name across reloads. To be safe, it also copies the state map rather than just its pointer so that remaining collisions disappear after one evaluation interval.	9 years ago
Fabian Reinartz	bfa8aaa017	Rename notification to notifier	9 years ago
beorn7	663a1550d0	Fix the instrumentation fixes	9 years ago
Tobias Schmidt	f1f8317fa5	Fix detection of flapping alerts Alerts in the resolve retention period must be transitioned to the active state again when their condition is met.	9 years ago
Björn Rabenstein	9ea3897ea7	Merge pull request #1354 from prometheus/beorn7/storage Rework the way to communicate backpressure (AKA suspended ingestion)	9 years ago
beorn7	ec08c9a391	Rework the way to communicate backpressure (AKA suspended ingestion) This gives up on the idea to communicate throuh the Append() call (by either not returning as it is now or returning an error as suggested/explored elsewhere). Here I have added a Throttled() call, which has the advantage that it can be called before a whole _batch_ of Append()'s. Scrapes will happen completely or not at all. Same for rule group evaluations. That's a highly desired behavior (as discussed elsewhere). The code is even simpler now as the whole ingestion buffer could be removed. Logging of throttled mode has been streamlined and will create at most one message per minute.	9 years ago
beorn7	a7408bfb47	Unify duration parsing It's actually happening in several places (and for flags, we use the standard Go time.Duration...). This at least reduces all our home-grown parsing to one place (in model).	9 years ago
Fabian Reinartz	a6935024e1	Remove old WITH clause in alert printing	9 years ago
Fabian Reinartz	b0adfea8d5	Fix swapped constants, improve instrumentation	9 years ago
Fabian Reinartz	a8c38c3ac5	Don't log rule evaluation failure on shutdown	9 years ago
Fabian Reinartz	6eee86dce8	Terminate rule groups during initial sleep When an evaluation group runs initially, it waits a deterministic amount of time. During that time it also has to accept a termination singnal so shutdown doesn't hang during the first evaluation iteration after a configuration reload. Fixes #1307	9 years ago
Fabian Reinartz	26eb3ac2f8	Don't skip recording rule errors	9 years ago
Fabian Reinartz	37d80c4b25	Fix premature rule evaluation This commit prevents rule evaluation from starting until after the storage is ready.	9 years ago
Fabian Reinartz	0cf3c6a9ef	Add comments, rename a method	9 years ago
Fabian Reinartz	bf6abac8f4	Send resolved notifications	9 years ago
Fabian Reinartz	f69e668fc4	Improve rules/ instrumentation This commit adds a counter for the total number of rule evaluations and standardizes the units to seconds.	9 years ago
Fabian Reinartz	62075aa037	Reduce noisy no-alertmanager warning	9 years ago
Fabian Reinartz	52e5224f5a	Refactor rules/ package	9 years ago
Fabian Reinartz	e4fabe135a	Set StartsAt to time of first firing state	9 years ago
Fabian Reinartz	7c90db22ed	Use annotation based alerts in rules/ This commit breaks the previously used alert format.	9 years ago
Fabian Reinartz	e114ce0ff7	Refactor notification handler	9 years ago
Fabian Reinartz	e3b6ec9784	Switch to common/log	9 years ago
Fabian Reinartz	171f50706a	Fix unkeyed field errors.	9 years ago
Brian Brazil	4d196fea6b	Merge pull request #1032 from prometheus/scalar-metric rules: Allow for setting labels on LHS on scalars	9 years ago
Brian Brazil	3bcdb2bbba	rules: Allow for setting labels on LHS on scalars	9 years ago
Julius Volz	995d3b831d	Fix most golint warnings. This is with `golint -min_confidence=0.5`. I left several lint warnings untouched because they were either incorrect or I felt it was better not to change them at the moment.	9 years ago
Fabian Reinartz	d6b8da8d43	Switch promql types to common/model	9 years ago
Brian Brazil	fdf0d0642e	Cast value to float, as that's what the console templates expect.	9 years ago
Fabian Reinartz	438e232c9b	Fix grouping of import blocks	9 years ago
Fabian Reinartz	306e8468a0	Switch from client_golang/model to common/model	9 years ago
Brian Brazil	e6a67476c2	rules: Allow recorded rules expressions to be scalars. This is useful if you want to build up a constant metric, such as a set of alert thresholds that vary by label value.	9 years ago
Fabian Reinartz	7a67472fc1	Resolve relative paths on configuration loading This moves the concern of resolving the files relative to the config file into the configuration loading itself. It also fixes #921 which did not load the cert and token files relatively.	9 years ago
Fabian Reinartz	feb8a03503	rules: load rule files relative to a base dir	10 years ago
Julius Volz	fcff35b43e	Consolidate external reachability flags into one. Besides fixing https://github.com/prometheus/prometheus/issues/805 by making the entire externally reachable server URL configurable, this adds tests for the "globalURL" template function and makes it easier to test other such functions in the future. This breaks the `web.Hostname` flag (and introduces `web.external-url`). This flag is likely only used by few users, so I hope that's justifiable. Fixes https://github.com/prometheus/prometheus/issues/805	10 years ago
Fabian Reinartz	f06cf664e1	rules: cleanup alerting test	10 years ago
Fabian Reinartz	9bd4f6d017	rules: preserve alert state across reloads.	10 years ago
Fabian Reinartz	4625485b84	rules: move rules.go contents to manager.go	10 years ago
Fabian Reinartz	749ae450c5	promql: add runbook to alert statement. This commit adds the RUNBOOK keyword to alert statements. The field is optional and expected to be a link.	10 years ago
Julius Volz	d868264bb8	Improve UI of /alerts page. Changes to the UI: - "Active Since" timestamps are now human-readable. - Alerting rules are now pretty-printed better. - Labels are no longer just strings, but alert bubbles (like we do on the status page for base labels). - Alert states and target health states are now capitalized in the presentation layer rather than at the source.	10 years ago
Fabian Reinartz	fe301d7946	promql: remove global flags	10 years ago
Fabian Reinartz	5e13880201	General cleanup of rules.	10 years ago
Fabian Reinartz	75c920c95e	Remove DotGraph method from Rule interface	10 years ago
Fabian Reinartz	83d07516e8	Remove EvalRaw methods from Rule interface	10 years ago
Fabian Reinartz	280d11dca8	main: exit on invalid rule files on startup.	10 years ago
Fabian Reinartz	0de6edbdfc	Move pkg/ to util/	10 years ago
Fabian Reinartz	02717e6fde	Remove generic set type	10 years ago
Fabian Reinartz	dbc0d30e3e	Move string functionality to pkg/strutil	10 years ago
Fabian Reinartz	f45a5cab60	Move templates package to pkg/template	10 years ago
Fabian Reinartz	c44ac7bc26	Load rule files from entire directories	10 years ago
Julius Volz	d7c015c149	Convert pathPrefix to not have trailing slash.	10 years ago
Julius Volz	ff53d10849	Fix double slash in GeneratorURL sent to alertmanager. Fixes https://github.com/prometheus/prometheus/issues/722	10 years ago
Julius Volz	267fd34156	Switch Prometheus to use github.com/prometheus/log. This change is conceptually very simple, although the diff is large. It switches logging from "github.com/golang/glog" to "github.com/prometheus/log", while not actually changing any log messages. V(1)-style logging has been changed to be log.Debug*().	10 years ago
Fabian Reinartz	e2ed921505	Merge branch 'master' into fabxc/servdisc	10 years ago
Mitsuhiro Tanda	3e914a8cb1	fix graph links with path prefix	10 years ago
Fabian Reinartz	bb540fd9fd	Implement config reloading on SIGHUP. With this commit, sending SIGHUP to the Prometheus process will reload and apply the configuration file. The different components attempt to handle failing changes gracefully.	10 years ago
Fabian Reinartz	fe935179cd	Stop routing rule statements through the engine.	10 years ago
Fabian Reinartz	8d7c479fed	Merge pull request #658 from prometheus/fabxc/pql/rules-manager Rename RuleManager to Manager, remove interface.	10 years ago
Fabian Reinartz	479891c9be	Rename RuleManager to Manager, remove interface. This commits renames the RuleManager to Manager as the package name is 'rules' now. The unused layer of abstraction of the RuleManager interface is removed.	10 years ago
Fabian Reinartz	25cdff3527	Remove `name` arg from `Parse*` functions, enhance parsing errors.	10 years ago
Fabian Reinartz	3ca11bcaf5	Switch Prometheus to promql package. This commit removes all functionality from rules/ that is now handled in promql/. All parts of Prometheus are changed to use the promql/ package.	10 years ago
Ceesjan Luiten	0e18784c64	Make all paths absolute to support proxies	10 years ago
Brian Brazil	941f585164	Avoid +InfYs and similar, just display +Inf.	10 years ago
beorn7	a075900f9a	Merge branch 'beorn7/persistence' into beorn7/ingestion-tweaks	10 years ago
Fabian Reinartz	624f27f4b6	Add ln, log2, log10 and exp functions to the query language.	10 years ago
Julius Volz	b2651027fc	Fix special value handling in division and modulo. This fixes https://github.com/prometheus/prometheus/issues/597	10 years ago
beorn7	be11cb2b07	Remove the sample ingestion channel. The one central sample ingestion channel has caused a variety of trouble. This commit removes it. Targets and rule evaluation call an Append method directly now. To incorporate multiple storage backends (like OpenTSDB), storage.Tee forks the Append into two different appenders. Note that the tsdb queue manager had its own queue anyway. It was a queue after a queue... Much queue, so overhead... Targets have their own little buffer (implemented as a channel) to avoid stalling during an http scrape. But a new scrape will only be started once the old one is fully ingested. The contraption of three pipelined ingesters was removed. A Target is an ingester itself now. Despite more logic in Target, things should be less confusing now. Also, remove lint and vet warnings in ast.go.	10 years ago
beorn7	13fcf1ddbc	Implement double-delta encoded chunks.	10 years ago
beorn7	9e85ab0eef	Apply the new signature/fingerprinting functions from client_golang. This requires the new version of client_golang (vendoring will follow in the next commit), which changes the fingerprinting for clientmodel.Metric.	10 years ago
Fabian Reinartz	182de6b99f	Fix unary +/- expressions. Unary expressions cause parsing errors if they are done in the lexer by tokenizing them into the number. This fix moves unary expressions to the parser.	10 years ago
Fabian Reinartz	6f754073d5	Add OR operation and vector matching options. This commits implements the OR operation between two vectors. Vector matching using the ON clause is added to limit the set of labels that define a match between two elements. Group modifiers (GROUP_LEFT/GROUP_RIGHT) to request many-to-one matching are added.	10 years ago
Julius Volz	0ac931aed1	Also support parsing float formats like "2.".	10 years ago
Julius Volz	c2ab54e9a6	Support scientific notation and special float values. This adds support for scientific notation in the expression language, as well as for all possible literal forms of +Inf/-Inf/NaN. TODO: Keep enough state in the parser/lexer to distinguish contexts in which "Inf", "NaN", etc. should be parsed as a number vs. parsed as a label name. Currently, foo{nan="bar"} would be a syntax error. However, that is an existing bug for all our reserved words. E.g. foo{sum="bar"} is a syntax error as well. This should be fixed separately.	10 years ago
beorn7	1a61bcae07	Fix plural of 'histogram'. Actually, 'histogram' is Ancient Greek and 3rd declension... ;-)	10 years ago
beorn7	17443d288b	Avoid copying of the COWMetric if we already have the metric available.	10 years ago
beorn7	9e7c3e3bcd	Add the histogram_quantile function. Since we are now getting really deep into floating point calculation, the tests had to take into account the precision loss. Since the rule tests are based on direct line matching in the output, implementing the "almost equal" semantics was pretty cumbersome, but here we are.	10 years ago
Julius Volz	42601acfde	Replace labelsToKey() with metric Fingerprint (fixes grouping bug).	10 years ago
Julius Volz	7fefccd929	Write() directly into hash and use model.SeparatorByte.	10 years ago
Julius Volz	645cf57bed	Fix aggregation grouping key calculation.	10 years ago
Julius Volz	15b2b5aa66	Add tests for invalid uses of "offset".	10 years ago
Julius Volz	67e20acc6c	Lower-case some package-internal names.	10 years ago
Julius Volz	72d7b325a1	Implement offset operator. This allows changing the time offset for individual instant and range vectors in a query. For example, this returns the value of `foo` 5 minutes in the past relative to the current query evaluation time: foo offset 5m Note that the `offset` modifier always needs to follow the selector immediately. I.e. the following would be correct: sum(foo offset 5m) // GOOD. While the following would be incorrect: sum(foo) offset 5m // INVALID. The same works for range vectors. This returns the 5-minutes-rate that `foo` had a week ago: rate(foo[5m] offset 1w) This change touches the following components: * Lexer/parser: additions to correctly parse the new `offset`/`OFFSET` keyword. * AST: vector and matrix nodes now have an additional `offset` field. This is used during their evaluation to adjust query and result times appropriately. * Query analyzer: now works on separate sets of ranges and instants per offset. Isolating different offsets from each other completely in this way keeps the preloading code relatively simple. No storage engine changes were needed by this change. The rules tests have been changed to not probe the internal implementation details of the query analyzer anymore (how many instants and ranges have been preloaded). This would also become too cumbersome to test with the new model, and measuring the result of the query should be sufficient. This fixes https://github.com/prometheus/prometheus/issues/529 This fixed https://github.com/prometheus/promdash/issues/201	10 years ago
Brian Brazil	60271d58bf	Change the 2nd argument of round to toNearest. This is more useful if you want get a multiple of 2 or 5, while still working for .001.	10 years ago
Julius Volz	82613527f3	Remove unnecessary float64() conversion in round().	10 years ago
Marko Mikulicic	8fdacbdf17	Add floor, ceil and round functions. Closes #402	10 years ago
Fabian Reinartz	fa1e90003b	Query timeout added. This is related to #454. Queries now timeout after a duration set by the -query.timeout flag. The TotalEvalTimer is now started/stopped inside any of the ast.Eval* functions.	10 years ago
Bjoern Rabenstein	26e22e6ad6	Fix rule manager shutdown.	10 years ago
Julius Volz	d4374a9265	More efficient JSON query result format. This depends on https://github.com/prometheus/client_golang/pull/51. For vectors, the result format looks like this: ```json { "version": 1, "type" : "vector", "value" : [ { "timestamp" : 1421765411.045, "value" : "65.475000", "metric" : { "quantile" : "0.5", "instance" : "http://localhost:9090/metrics", "job" : "prometheus", "__name__" : "http_request_duration_microseconds", "handler" : "/static/", "method" : "get", "code" : "304" } }, { "timestamp" : 1421765411.045, "value" : "5826.339000", "metric" : { "quantile" : "0.9", "instance" : "http://localhost:9090/metrics", "job" : "prometheus", "__name__" : "http_request_duration_microseconds", "handler" : "prometheus", "method" : "get", "code" : "200" } }, /* ... / ] } ``` For matrices, it looks like this: ```json { "version": 1, "type" : "matrix", "value" : [ { "metric" : { "quantile" : "0.99", "instance" : "http://localhost:9090/metrics", "job" : "prometheus", "__name__" : "http_request_duration_microseconds", "handler" : "/static/", "method" : "get", "code" : "200" }, "values" : [ [ 1421765547.659, "29162.953000" ], [ 1421765548.659, "29162.953000" ], [ 1421765549.659, "29162.953000" ], / ... */ ] } ] } ```	10 years ago
Brian Brazil	a31730e88b	Make 2nd arg to delta optional. Add a deriv() function. The 2nd isCounter argument to delta is ugly, make it optional as the first step of deprecating it. This will makes delta only ever applied to gauges. Add a deriv function to calculate the least squares slope of a gauge. This is more useful for prediction than delta, as it isn't as heavily influenced by outliers at the boundaries.	10 years ago
Bjoern Rabenstein	5859b74f1b	Clean up license issues. - Move CONTRIBUTORS.md to the more common AUTHORS. - Added the required NOTICE file. - Changed "Prometheus Team" to "The Prometheus Authors". - Reverted the erroneous changes to the Apache License.	10 years ago
Bjoern Rabenstein	b09453af1d	Adjust to new client_golang API.	10 years ago
Julius Volz	bb1e49383e	Log rule evalation errors.	10 years ago
Julius Volz	d6b9e97655	Remove extraction.Result type, simplify code.	10 years ago
Julius Volz	9a4ca68a61	Add metrics for rule evaluation failures. Fixes https://github.com/prometheus/prometheus/issues/417	10 years ago
Brian Brazil	ffa2e73803	Fix regression from `5e8d57bec1` 0 is a false value, so shortcutting no longer works. Update other places in the code that assumed graph was the default.	10 years ago
Julius Volz	cc27fb8aab	Rename remaining all-caps constants in AST layer. Change-Id: Ibe97e30981969056ffcdb89e63c1468ea1ffa140	10 years ago
Julius Volz	895523ad14	Include necessary Makefile.INCLUDE from rules/Makefile. Change-Id: I077d018dbe4093cd40ddf38d66a996df222bf5e4	10 years ago
Julius Volz	2ade9d40cf	Clarify why we need int constants for expression types. Change-Id: I053fc5d32c118dbdb204dc8193337f981aff796e	10 years ago
Julius Volz	00a2a93a05	Add regression tests for metrics mutations in AST. It turned out in the end, that only drop_common_metrics() produced any erroneous output in the old system. The second expression in the test ("sum(testmetric) keeping_extra") already worked in the old code, but why not keep it in... The way to test ranged evaluations is a bit clumsy so far, so I want to build a nicer test framework in the end, where all the test cases can be specified as text files which specify desired inputs, outputs, query step widths, etc. Change-Id: I821859789e69b8232bededf670a1b76e9e8c8ca4	10 years ago
Julius Volz	c9618d11e8	Introduce copy-on-write for metrics in AST. This depends on changes in: https://github.com/prometheus/client_golang/tree/cow-metrics. Change-Id: I80b94833a60ddf954c7cd92fd2cfbebd8dd46142	10 years ago
Bjoern Rabenstein	b1e4956142	Apply a giant code cleanup. Essentially: - Remove unused code. - Make it 'go vet' clean. The only remaining warnings are in generated code. - Make it 'golint' clean. The only remaining warnings are in gerenated code. - Smoothed out same minor things. Change-Id: I3fe5c1fbead27b0e7a9c247fee2f5a45bc2d42c6	10 years ago
Bjoern Rabenstein	fee88a7a77	Remove the remaining races, new and old. Also, resolve a few other TODOs. Change-Id: Icb39b5a5e8ca22ebcb48771cd8951c5d9e112691	10 years ago
Bjoern Rabenstein	7d11019aa2	Squash a few trivial TODOs. - Delete unneeded file view_adapter.go. - Assessed that we still need the fingerprints in nodes (to create iterators). - Turned numMemChunkDescs into a metric. Change-Id: I29be963c795a075ec00c095f76bf26405535609d	10 years ago
Julius Volz	6eecee55b7	Fix acronym caps in GeneratorURL. Change-Id: Ib18c1f617dcde1039e848059545a6d8831d9bf66	10 years ago
Bjoern Rabenstein	0ae1d8889a	Fix tests after merge. Change-Id: Ia90da9a3e48ed780ec38c4a6a1fd9ea34e7f6a58	10 years ago
Julius Volz	b7bf11230a	Add absent() function. A common problem in Prometheus alerting is to detect when no timeseries exist for a given metric name and label combination. Unfortunately, Prometheus alert expressions need to be of vector type, and "count(nonexistent_metric)" results in an empty vector, yielding no output vector elements to base an alert on. The newly introduced absent() function solves this issue: ALERT FooAbsent IF absent(foo{job="myjob"}) [...] absent() has the following behavior: - if the vector passed to it has any elements, it returns an empty vector. - if the vector passed to it has no elements, it returns a 1-element vector with the value 1. In the second case, absent() tries to be smart about deriving labels of the 1-element output vector from the input vector: absent(nonexistent{job="myjob"}) => {job="myjob"} absent(nonexistent{job="myjob",instance=~".*"}) => {job="myjob"} absent(sum(nonexistent{job="myjob"})) => {} That is, if the passed vector is a literal vector selector, it takes all "=" label matchers as the basis for the output labels, but ignores all non-equals or regex matchers. Also, if the passed vector results from a non-selector expression, no labels can be derived. Change-Id: I948505a1488d50265ab5692a3286bd7c8c70cd78	10 years ago
Julius Volz	3d47f94149	Drop metric names after transformations. After many transformations, it doesn't make sense to keep the metric names, since the result of the transformation is no longer that metric. This drops the metric name after such transformations and makes the web UI deal well with missing metric names. This depends on the current branch on the following things: - prometheus/client_golang needs to be at `e237cf15c6` in branch "julius/int-fingerprints" (to be merged with new storage) - prometheus/promdash needs to be at `dd7691c9c2` Change-Id: Ib3c8cad8d647d9854e8c653c424b8c235ccc231d	10 years ago
Bjoern Rabenstein	14bda4180c	Changes after pair code review. Change-Id: Ib72d40f8e9027818cfbbd32a7a7201eebda07455	10 years ago
Bjoern Rabenstein	006b5517e2	Simplify makefiles. This removes the dependancy on C leveldb and snappy. It also takes care of fewer dependencies as they would anyway not work on any non-Debian, non-Brew system. Change-Id: Ia70dce1ba8a816a003587927e0b3a3f8ad2fd28c	10 years ago
Bjoern Rabenstein	74c143c4c9	Improve scraper shutdown time. - Stop target pools in parallel. - Stop individual scrapers in goroutines, too. - Timing tweaks. Change-Id: I9dff1ee18616694f14b04408eaf1625d0f989696	10 years ago
Julius Volz	0712d738d1	Allow alternative "by"-clause position in grammar. In addition to the existing by-clause syntax: sum(<expression>) by (<labels>) [keeping_extra] ...this allows the following new syntax: sum by (<labels>) [keeping_extra] (<expression>) Both orderings may be used in a single expression. It is up to the users to establish guidelines around their usage. Change-Id: Iba10c9cc5fb6ac62edfcf246d281473e82467992	10 years ago
Julius Volz	0e48c18bbf	Allow omitting the metric name in queries. This allows the following expression syntaxes for selecting timeseries: foo (already valid before) foo{} (already valid before) {job="prometheus"} (new, select all timeseries for job "prometheus") Omitting both the metric name and any label matchers ("" or "{}") will still yield a syntax error. To get all timeseries, you could do: {__name__=~"."} or, without relying on knowledge about __metric__: {job=~"."} Change-Id: Ifee000b9ac0184ef6ced18411069c7f2699a2dda	10 years ago
Bjoern Rabenstein	096fa0f8b2	Squash a number of TODOs. - Staleness delta is no a proper function parameter and not replicated from package ast. - Named type 'chunks' replaced by explicit '[]chunk' to avoid confusion. - For the same reason, replaced 'chunkDescs' by '[]*chunkDescs'. - Verified that math.Modf is not a speed enhancement over conversion (actually 5x slower). - Renamed firstTimeField, lastTimeField into chunkFirstTime and chunkLastTime. - Verified unpin() is sufficiently goroutine-safe. - Decided not to update archivedFingerprintToTimeRange upon series truncation and added a rationale why. Change-Id: I863b8d785e5ad9f71eb63e229845eacf1bed8534	10 years ago
Bjoern Rabenstein	b3ed9aa7a2	Clean up start-up and shut-down. Change-Id: Idff4bbb0a15a9f879bfbb3da5b1025179cab5e2c	10 years ago
Bjoern Rabenstein	38fc24d0ed	Fix targetpool_test.go and other tests. Change-Id: I91a4dd1d39e01f174e1aaae653ce1ed7aecaa624	10 years ago
Julius Volz	7f5d3c2c29	Fix and improve the fp locker. Benchmark: $ go test -bench 'Fingerprint' -test.run 'Fingerprint' -test.cpu=1,2,4 OLD BenchmarkFingerprintLockerParallel 500000 3618 ns/op BenchmarkFingerprintLockerParallel-2 100000 12257 ns/op BenchmarkFingerprintLockerParallel-4 500000 10164 ns/op BenchmarkFingerprintLockerSerial 10000000 283 ns/op BenchmarkFingerprintLockerSerial-2 10000000 284 ns/op BenchmarkFingerprintLockerSerial-4 10000000 288 ns/op NEW BenchmarkFingerprintLockerParallel 1000000 1018 ns/op BenchmarkFingerprintLockerParallel-2 1000000 1164 ns/op BenchmarkFingerprintLockerParallel-4 2000000 910 ns/op BenchmarkFingerprintLockerSerial 50000000 56.0 ns/op BenchmarkFingerprintLockerSerial-2 50000000 47.9 ns/op BenchmarkFingerprintLockerSerial-4 50000000 54.5 ns/op Change-Id: I3c65a43822840e7e64c3c3cfe759e1de51272581	10 years ago
Julius Volz	358f97791d	Minor cleanups. Change-Id: Ia8685d8439a421fe2143d9ec7120d5bb5ab88d78	10 years ago
Bjoern Rabenstein	f5f9f3514a	Major code cleanup. - Make it go-vet and golint clean. - Add comments, TODOs, etc. Change-Id: If1392d96f3d5b4cdde597b10c8dff1769fcfabe2	10 years ago
Julius Volz	e7ed39c9a6	Initial experimental snapshot of next-gen storage. Change-Id: Ifb8709960dbedd1d9f5efd88cdd359ee9fa9d26d	10 years ago
Julius Volz	85497e3f38	Add function to drop common labels in a vector. This fixes https://github.com/prometheus/prometheus/issues/384. Change-Id: I2973c4baeb8a4618ec3875fb11c6fcf5d111784b	10 years ago
Julius Volz	3fdb74e571	Add more topk() / bottomk() tests. Test what happens if k > number of input elements. Change-Id: Ie724b850939e297ebf085f0a5a3522e9cfcc6534	10 years ago
Julius Volz	c582ae73c2	Implement topk() and bottomk() functions. To achieve O(log n * k) runtime, this uses a heap to track the current bottom-k or top-k elements while iterating over the full set of available elements. It would be possible to reuse more code between topk and bottomk, but I decided for some more duplication for the sake of clarity. This fixes https://github.com/prometheus/prometheus/issues/399 Change-Id: I7487ddaadbe7acb22ca2cf2283ba6e7915f2b336	10 years ago
Bjoern Rabenstein	1909686789	Make metrics exported by the Prometheus server itself more consistent. - Always spell out the time unit (e.g. milliseconds instead of ms). - Remove "_total" from the names of metrics that are not counters. - Make use of the "Namespace" and "Subsystem" fields in the options. - Removed the "capacity" facet from all metrics about channels/queues. These are all fixed via command line flags and will never change during the runtime of a process. Also, they should not be part of the same metric family. I have added separate metrics for the capacity of queues as convenience. (They will never change and are only set once.) - I left "metric_disk_latency_microseconds" unchanged, although that metric measures the latency of the storage device, even if it is not a spinning disk. "SSD" is read by many as "solid state disk", so it's not too far off. (It should be "solid state drive", of course, but "metric_drive_latency_microseconds" is probably confusing.) - Brian suggested to not mix "failure" and "success" outcome in the same metric family (distinguished by labels). For now, I left it as it is. We are touching some bigger issue here, especially as other parts in the Prometheus ecosystem are following the same principle. We still need to come to terms here and then change things consistently everywhere. Change-Id: If799458b450d18f78500f05990301c12525197d3	10 years ago
Julius Volz	00b9489f1c	Fix time() behavior. time() should return the timestamp for which the query is executed, not the actual current time. Change-Id: I430a45cabad7785cd58f95b1028a71dff4c87710	10 years ago
Julius Volz	c5984f1818	Add abs() and over-time aggregation functions. This implements aggregation functions over time as request in https://github.com/prometheus/prometheus/issues/383. Change-Id: Ifd69b850de8cfdf6e7a6c0e042056fa4c672410e	10 years ago

... 3 4 5 6 7 ...

586 Commits (0cc99e677ad3da2cf00599cb0e6c272ab58688f1)