prometheus

Commit Graph

Author	SHA1	Message	Date
Tom Wilkie	48e39068bd	Don't allocate a mergeSeries if there is only one series to merge.	7 years ago
Bryan Boreham	8a4535e6ad	Re-use timer instead of creating new ones on every sample The docs for `time.After()` note that "The underlying Timer is not recovered by the garbage collector until the timer fires".	7 years ago
Tom Wilkie	f2c5399e39	Merge pull request #3561 from twiedenbein/master fixed bug with initialization of queueconfig	7 years ago
Shubheksha Jalan	0471e64ad1	Use shared types from the `common` repo (#3674 ) * refactor: use shared types from common repo, remove util/config * vendor: add common/config * fix nit	7 years ago
Shubheksha Jalan	ec94df49d4	Refactor SD configuration to remove `config` dependency (#3629 ) * refactor: move targetGroup struct and CheckOverflow() to their own package * refactor: move auth and security related structs to a utility package, fix import error in utility package * refactor: Azure SD, remove SD struct from config * refactor: DNS SD, remove SD struct from config into dns package * refactor: ec2 SD, move SD struct from config into the ec2 package * refactor: file SD, move SD struct from config to file discovery package * refactor: gce, move SD struct from config to gce discovery package * refactor: move HTTPClientConfig and URL into util/config, fix import error in httputil * refactor: consul, move SD struct from config into consul discovery package * refactor: marathon, move SD struct from config into marathon discovery package * refactor: triton, move SD struct from config to triton discovery package, fix test * refactor: zookeeper, move SD structs from config to zookeeper discovery package * refactor: openstack, remove SD struct from config, move into openstack discovery package * refactor: kubernetes, move SD struct from config into kubernetes discovery package * refactor: notifier, use targetgroup package instead of config * refactor: tests for file, marathon, triton SD - use targetgroup package instead of config.TargetGroup * refactor: retrieval, use targetgroup package instead of config.TargetGroup * refactor: storage, use config util package * refactor: discovery manager, use targetgroup package instead of config.TargetGroup * refactor: use HTTPClient and TLS config from configUtil instead of config * refactor: tests, use targetgroup package instead of config.TargetGroup * refactor: fix tagetgroup.Group pointers that were removed by mistake * refactor: openstack, kubernetes: drop prefixes * refactor: remove import aliases forced due to vscode bug * refactor: move main SD struct out of config into discovery/config * refactor: rename configUtil to config_util * refactor: rename yamlUtil to yaml_config * refactor: kubernetes, remove prefixes * refactor: move the TargetGroup package to discovery/ * refactor: fix order of imports	7 years ago
Ed Schouten	bb724f1bef	Deprecate DeduplicateSeriesSet() in favor of NewMergeSeriesSet(). Federation makes use of dedupedSeriesSet to merge SeriesSets for every query into one output stream. If many match[] arguments are provided, many dedupedSeriesSet objects will get chained. This has the downside of causing a potential O(nk) running time, where n is the number of series and k the number of match[] arguments. In the mean time, the storage package provides a mergeSeriesSet that accomplishes the same with an O(nlog(k)) running time by making use of a binary heap. Let's just get rid of dedupedSeriesSet and change all existing callers to use mergeSeriesSet.	7 years ago
Tom Wiedenbein	937ac8c060	fixed bug with initialization of queueconfig QueueConfigs would only ever initialize to the default settings, and would not pick up their respective values from YAML.	7 years ago
Fabian Reinartz	83cd270ea4	*: adapt to storage interface changes	7 years ago
Tobias Schmidt	7098c56474	Add remote read filter option For special remote read endpoints which have only data for specific queries, it is desired to limit the number of queries sent to the configured remote read endpoint to reduce latency and performance overhead.	7 years ago
Tobias Schmidt	434f0374f7	Refactor remote storage querier handling * Decouple remote client from ReadRecent feature. * Separate remote read filter into a small, testable function. * Use storage.Queryable interface to compose independent functionalities.	7 years ago
Tobias Schmidt	9b0091d487	Add storage.Queryable and storage.QueryableFunc In order to compose different querier implementations more easily, this change introduces a separate storage.Queryable interface grouping the query (Querier) function of the storage. Furthermore, it adds a QueryableFunc type to ease writing very simple queryable implementations.	7 years ago
Julius Volz	9f10c63cff	Fix remote read labelset corruption (#3456 ) The labelsets returned from remote read are mutated in higher levels (like seriesFilter.Labels()) and since the concreteSeriesSet didn't return a copy, the external mutation affected the labelset in the concreteSeries itself. This resulted in bizarre bugs where local and remote series would show with identical label sets in the UI, but not be deduplicated, since internally, a series might come to look like: {__name__="node_load5", instance="192.168.1.202:12090", job="node_exporter", node="odroid", node="odroid"} (note the repetition of the last label)	7 years ago
Krasi Georgiev	5d8f93a22a	now using only github.com/gogo/protobuf bumped all grpc-gateway packages to v1.2.2 updated and run the denproto.sh script	7 years ago
Fabian Reinartz	30e777d10d	tsdb: default too small max block duration	7 years ago
Tom Wilkie	48a7a00a38	Fast path the merge querier (#3358 ) * Fast path the merge querier such that it is completely removed from query path when there is no remote storage. * Add NoopQuerier * Add copyright notice. * Avoid global, use a function.	7 years ago
Tom Wilkie	0e572686db	Revert "Bypass the fanout storage merging if no remote storage is configured."	7 years ago
Tom Wilkie	1af3ef431d	s/TestRemoveLabels/TestSeriesSetFilter/	7 years ago
Tom Wilkie	9c3c98e8de	Revert "Port 'Don't disable HTTP keep-alives for remote storage connections.' to 2.0 (see #3173 )" This reverts commit `0997191b18`.	7 years ago
Tom Wilkie	746752b946	Merge external labels in order.	7 years ago
Tom Wilkie	6e4d4ea402	Initialise some counters in remote storage API.	7 years ago
Tom Wilkie	2ae04d0e79	Add license header.	7 years ago
Tom Wilkie	e8c264e47a	Add comment.	7 years ago
Tom Wilkie	ee011d906d	Port remote read server to 2.0.	7 years ago
Bryan Boreham	0997191b18	Port 'Don't disable HTTP keep-alives for remote storage connections.' to 2.0 (see #3173 ) Removes configurability introduced in #3160 in favour of hard-coding, per advice from @brian-brazil.	7 years ago
Tom Wilkie	56820726fa	Move a couple of the encoding/decoding functions into codec.go	7 years ago
Conor Broderick	08b7328669	Port Metric name validation to 2.0 (see #2975 )	7 years ago
Tom Wilkie	8fe0212ff7	Port 'Make queue manager configurable.' to 2.0, see #2991	7 years ago
Tom Wilkie	3760f56c0c	remote: Expose ClientConfig type (see #3165 )	7 years ago
Tom Wilkie	16f71a7723	Port codec.go over form 1.8 branch.	7 years ago
Fabian Reinartz	e53040e2ac	Merge pull request #3339 from tomwilkie/3065-remote-read-bypass Bypass the fanout storage merging if no remote storage is configured.	7 years ago
Fabian Reinartz	bf56ad4233	Merge branch 'master' into master	7 years ago
Paul Gier	c4c3205d76	storage/tsdb: check that max block duration is larger than min If the user accidentally sets the max block duration smaller than the min, the current error is not informative. This change just performs the check earlier and improves the error message.	7 years ago
Fabian Reinartz	ce63a5a855	Merge pull request #3352 from prometheus/rc2 Cut v2.0.0-rc.2	7 years ago
Thibault Chataigner	fc4406201e	Tsdb StartTime : Use a simplier way to compute StartTime	7 years ago
Julius Volz	099df0c5f0	Migrate "golang.org/x/net/context" -> "context" (#3333 ) In some places, where ctxhttp or gRPC are concerned, we still need to use the old contexts.	7 years ago
Tom Wilkie	4bbef0ec30	Bypass the fanout storage merging if no remote storage is configured.	7 years ago
Fabian Reinartz	a57ea79660	Close index reader properly	7 years ago
Julius Volz	c3d6abc8e6	Fix some lint errors (#3334 ) I left the promql ones and some others untouched as I remember that @fabxc prefers them that way.	7 years ago
Julius Volz	2846d62573	Fix staticcheck issue in test (#3331 ) staticcheck fails with: storage/remote/read_test.go:199:27: do not pass a nil Context, even if a function permits it; pass context.TODO if you are unsure about which Context to use (SA1012)	7 years ago
Brian Brazil	4a50f547c8	removeLabels needs a pointer to work. (#3326 )	7 years ago
Thibault Chataigner	bf4a279a91	Remote storage reads based on oldest timestamp in primary storage (#3129 ) Currently all read queries are simply pushed to remote read clients. This is fine, except for remote storage for wich it unefficient and make query slower even if remote read is unnecessary. So we need instead to compare the oldest timestamp in primary/local storage with the query range lower boundary. If the oldest timestamp is older than the mint parameter, then there is no need for remote read. This is an optionnal behavior per remote read client. Signed-off-by: Thibault Chataigner <t.chataigner@criteo.com>	7 years ago
Julius Volz	9ef8518b37	Remove "package remote" garbage from license headers (#3304 )	7 years ago
Tobias Schmidt	721050c6cb	Update prometheus/tsdb dependency	7 years ago
Julius Volz	33c1171b9c	Don't add anchoring to exported `Value` matcher field Instead, just make the anchoring part of the internal regex. This helps because some users will want to read back the `Value` field and expect it to be the same as the input value (e.g. some tests in Cortex), or use the value in another context which is already expected to add its own anchoring, leading to superfluous double anchoring (such as when we translate matchers into remote read request matchers).	7 years ago
Brian Brazil	73dc96e7f5	Fix leak of ticker in remote storage queue manager.	7 years ago
Brian Brazil	ee88f0d222	Ensure all values are used or _	7 years ago
Brian Brazil	37ec2d5283	Fix off by one error in concreteSeriesSet (#3262 )	7 years ago
Marc Sluiter	6a633eece1	Added go-conntrack for monitoring http connections (#3241 ) Added metrics for in- and outgoing traffic with go-conntrack.	7 years ago
Julius Volz	f7e8348a88	Re-add contexts to storage.Storage.Querier() (#3230 ) * Re-add contexts to storage.Storage.Querier() These are needed when replacing the storage by a multi-tenant implementation where the tenant is stored in the context. The 1.x query interfaces already had contexts, but they got lost in 2.x. * Convert promql.Engine to use native contexts	7 years ago
Fabian Reinartz	7b02bfee0a	web: start web handler while TSDB is starting up	7 years ago
Fabian Reinartz	d21f149745	*: migrate to go-kit/log	7 years ago
Fabian Reinartz	0efecea6d4	Adapt storage APIs to uint64 references	7 years ago
Fabian Reinartz	0c81d5f719	storage: instantiate correct block ranges	7 years ago
Fabian Reinartz	2037778d14	vendor: update TSDB	7 years ago
Tom Wilkie	b11bc8ae24	Fix some comments.	7 years ago
Tom Wilkie	ec999ff397	Prevent number of remote write shards from going negative. This can happen in the situation where the system scales up the number of shards massively (to deal with some backlog), then scales it down again as the number of samples sent during the time period is less than the number received.	7 years ago
Tom Wilkie	a09acdcc5b	Make concreteSeriersIterator behave.	7 years ago
Tom Wilkie	994a7f27d6	Propagate errors through mergeSeriesSet correctly.	7 years ago
Tom Wilkie	2e0d8487e3	Return zeros if At() is called after Next() returns false.	7 years ago
Tom Wilkie	014bd31a86	Remove unnecessary whitespace changes, add comment.	7 years ago
Tom Wilkie	98ac07f86a	Add unit test for the merging on the read path.	7 years ago
Tom Wilkie	b568ace7ce	Move protos to ./prompb	7 years ago
Tom Wilkie	96e25adc8d	Introduce 'primary' storage in fanout, and have Add return the ref from the primary. Also, ensure all append batches are rolled back when a commit or rollback fails.	7 years ago
Tom Wilkie	db8128ceeb	Add label set as first parameter to AddFast, ingored by TSDB adapter.	7 years ago
Tom Wilkie	2dda5775e3	Initial port of remote storage to v2.	7 years ago
Fabian Reinartz	16464c3a33	Merge pull request #2910 from prometheus/adminapi Admin API	7 years ago
Fabian Reinartz	ccf9e62972	*: add admin grpc API	7 years ago
Goutham Veeramachaneni	243419c007	Return tsdb.ErrOutOfBounds as storage.ErrOutOfBounds Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	3069bd3996	Handle scrapes with OutOfBounds metrics better fixes #2894 Signed-off-by: Goutham Veeramachaneni <goutham@boomerangcommerce.com>	8 years ago
Goutham Veeramachaneni	d407bd150c	Consolidate the duration params in CLI * All CLI params moved to model.Duration Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Goutham Veeramachaneni	baf5b0f0fc	Fix error where we look into the future. (#2829 ) * Fix error where we look into the future. So currently we are adding values that are in the future for an older timestamp. For example, if we have [(1, 1), (150, 2)] we will end up showing [(1, 1), (2,2)]. Further it is not advisable to call .At() after Next() returns false. Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in> * Retuen early if done Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in> * Handle Seek() where we reach the end of iterator Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in> * Simplify code Signed-off-by: Goutham Veeramachaneni <cs14btech11014@iith.ac.in>	8 years ago
Brian Brazil	c02c25d5ba	Allow peeking back further in buffer.	8 years ago
Fabian Reinartz	d289dc55c3	storage: update TSDB	8 years ago
Fabian Reinartz	9b175d48cb	Add flag to disable TSDB lock file	8 years ago
Fabian Reinartz	0f3110487d	Merge remote-tracking branch 'origin/dev-2.0' into dev-2.0	8 years ago
Fabian Reinartz	37deb21c45	vendor: remove unused dependency and last ref to fabxc/tsdb	8 years ago
Brian Brazil	5c9a6ce747	Add license to files. This should fix CI for dev-2.0.	8 years ago
Fabian Reinartz	8ffc851147	Merge branch 'master' into dev-2.0	8 years ago
Fabian Reinartz	cfb2a7f1d5	vendor: sync organisation migration of tsdb	8 years ago
Fabian Reinartz	bbcf20ba01	web: deduplicate series in federation	8 years ago
Fabian Reinartz	4e41987bcb	storage: add deduplication function This adds a function to deduplicate two series sets given that duplicate series have equivalent data points.	8 years ago
Björn Rabenstein	50e4f49b7e	Merge pull request #2561 from prometheus/beorn7/storage2 storage: Evict unused chunk.Descs in crash recovery	8 years ago
beorn7	08fc6cbd39	storage: Evict unused chunk.Descs in crash recovery This is in line with the v1.5 change in paradigm to not keep chunk.Descs without chunks around after a series maintenance. It's mainly motivated by avoiding excessive amounts of RAM usage during crash recovery. The code avoids to create memory time series with zero chunk.Descs as that is prone to trigger weird effects. (Series maintenance would archive series with zero chunk.Descs, but we cannot do that here because the archive indices still have to be checked.)	8 years ago
Björn Rabenstein	1c6240fc40	Merge pull request #2559 from prometheus/beorn7/storage storage: Replace fpIter by sortedFPs	8 years ago
beorn7	d284ffab03	storage: Replace fpIter by sortedFPs The fpIter was kind of cumbersome to use and required a lock for each iteration (which wasn't even needed for the iteration at startup after loading the checkpoint). The new implementation here has an obvious penalty in memory, but it's only 8 byte per series, so 80MiB for a beefy server with 10M memory time series (which would probably need ~100GiB RAM, so the memory penalty is only 0.1% of the total memory need). The big advantage is that now series maintenance happens in order, which leads to the time between two maintenances of the same series being less random. Ideally, after each maintenance, the next maintenance would tackle the series with the largest number of non-persisted chunks. That would be quite an effort to find out or track, but with the approach here, the next maintenance will tackle the series whose previous maintenance is longest ago, which is a good approximation. While this commit won't change the _average_ number of chunks persisted per maintenance, it will reduce the mean time a given chunk has to wait for its persistence and thus reduce the steady-state number of chunks waiting for persistence. Also, the map iteration in Go is non-deterministic but not truly random. In practice, the iteration appears to be somewhat "bucketed". You can often observe a bunch of series with similar duration since their last maintenance, i.e. you see batches of series with similar number of chunks persisted per maintenance. If that batch is relatively young, a whole lot of series are maintained with very few chunks to persist. (See screenshot in PR for a better explanation.)	8 years ago
Tobias Schmidt	eac36d123e	Fix unstable fanin test (#2558 )	8 years ago
Julius Volz	5a896033e3	Add remote read external label handling (#2555 ) * Add remote read external label handling This implements rule 1 and 2 from https://docs.google.com/document/d/188YauRgfF0J4CYMigLsVNN34V_kUwKnApBs2dQMfBbs/edit * Use more descriptive example labels in read test * Add comment for querier.addExternalLabels() * Make argument naming in removeLabels() more generic	8 years ago
Björn Rabenstein	e63d079b59	Merge pull request #2527 from prometheus/beorn7/storage storage: Evict chunks and calculate persistence pressure...	8 years ago
Julius Volz	b5b0e00923	Merge pull request #2499 from prometheus/remote-read Remote Read	8 years ago
beorn7	434ab2a6a3	storage: Evict chunks and calculate persistence pressure based on target heap size This is a fairly easy attempt to dynamically evict chunks based on the heap size. A target heap size has to be set as a command line flage, so that users can essentially say "utilize 4GiB of RAM, and please don't OOM". The -storage.local.max-chunks-to-persist and -storage.local.memory-chunks flags are deprecated by this change. Backwards compatibility is provided by ignoring -storage.local.max-chunks-to-persist and use -storage.local.memory-chunks to set the new -storage.local.target-heap-size to a reasonable (and conservative) value (both with a warning). This also makes the metrics intstrumentation more consistent (in naming and implementation) and cleans up a few quirks in the tests. Answers to anticipated comments: There is a chance that Go 1.9 will allow programs better control over the Go memory management. I don't expect those changes to be in contradiction with the approach here, but I do expect them to complement them and allow them to be more precise and controlled. In any case, once those Go changes are available, this code has to be revisted. One might be tempted to let the user specify an estimated value for the RSS usage, and then internall set a target heap size of a certain fraction of that. (In my experience, 2/3 is a fairly safe bet.) However, investigations have shown that RSS size and its relation to the heap size is really really complicated. It depends on so many factors that I wouldn't even start listing them in a commit description. It depends on many circumstances and not at least on the risk trade-off of each individual user between RAM utilization and probability of OOMing during a RAM usage peak. To not add even more to the confusion, we need to stick to the well-defined number we also use in the targeting here, the sum of the sizes of heap objects.	8 years ago
beorn7	96a303b348	storage: Use staleness delta as head chunk timeout Currently, if a series stops to exist, its head chunk will be kept open for an hour. That prevents it from being persisted. Which prevents it from being evicted. Which prevents the series from being archived. Most of the time, once no sample has been added to a series within the staleness limit, we can be pretty confident that this series will not receive samples anymore. The whole chain as described above can be started after 5m instead of 1h. In the relaxed case, this doesn't change a lot as the head chunk timeout is only checked during series maintenance, and usually, a series is only maintained every six hours. However, there is the typical scenario where a large service is deployed, the deoply turns out to be bad, and then it is deployed again within minutes, and quite quickly the number of time series has tripled. That's the point where the Prometheus server is stressed and switches (rightfully) into rushed mode. In that mode, time series are processed as quickly as possible, but all of that is in vein if all of those recently ended time series cannot be persisted yet for another hour. In that scenario, this change will help most, and it's exactly the scenario where help is most desperately needed.	8 years ago
Julius Volz	3f23aa2cc7	Add headers to indicate remote read/write version Also add Content-Type header.	8 years ago
Julius Volz	8fda83ea12	Make rules only read local data	8 years ago
Julius Volz	94acd3f1d8	Add fanin tests and fix uncovered bugs	8 years ago
Julius Volz	9b33cfc457	Fix/unify context-based remote storage timeouts	8 years ago
Julius Volz	815762a4ad	Move retrieval.NewHTTPClient -> httputil.NewClientFromConfig	8 years ago
Fabian Reinartz	397f001ac5	Merge branch 'master' into dev-2.0	8 years ago
Julius Volz	eb14678a25	Make remote read/write use config.HTTPClientConfig	8 years ago
Julius Volz	406b65d0dc	Rename remote.Storage to remote.Writer	8 years ago
Julius Volz	02395a224d	[WIP] Remote Read	8 years ago
Julius Volz	40e41a4776	Merge pull request #2494 from tomwilkie/remote-write-sharding Dynamically reshard the QueueManager based on observed load.	8 years ago
Fabian Reinartz	b586781283	*: update tsdb vendoring and add retention flag	8 years ago
beorn7	48d221c11e	storage: Fix typo in comment	8 years ago
Fabian Reinartz	0ecd205794	promql: Use buffer pool for matrix allocations	8 years ago
Tom Wilkie	75bb0f3253	Review feedback	8 years ago
Tom Wilkie	77cce900b8	Fix tests	8 years ago
Tom Wilkie	b48799a01e	Add license stanza	8 years ago
Tom Wilkie	9d22f030cf	Dynamically reshard the QueueManager based on observed load.	8 years ago
Fabian Reinartz	8a8eb12985	storage/tsdb: don't use partitioned DB.	8 years ago
Fabian Reinartz	9eb1d6c927	remote: take code from master	8 years ago
Fabian Reinartz	9304179ef7	Merge branch 'master' into dev-2.0	8 years ago
Fabian Reinartz	4397b4d508	*: pass Prometheus registry into storage	8 years ago
Tom Wilkie	1ab893c6ec	Limit 'discarding sample' logs to 1 every 10s (#2446 ) * Limit 'discarding sample' logs to 1 every 10s * Include the vendored library * Review feedback	8 years ago
Julius Volz	2f39dbc8b3	Rename StorageQueueManager -> QueueManager	8 years ago
Julius Volz	e9476b35d5	Re-add multiple remote writers Each remote write endpoint gets its own set of relabeling rules. This is based on the (yet-to-be-merged) https://github.com/prometheus/prometheus/pull/2419, which removes legacy remote write implementations.	8 years ago
Björn Rabenstein	089dc1076b	Merge pull request #2435 from jmeulemans/open-chunks-gauge Adding gauge for number of open head chunks.	8 years ago
Jeremy Meulemans	025c828976	Changed to open_head_chunks to address review. Now incrementing numHeadChunks directly.	8 years ago
Jeremy Meulemans	074050b8c0	Updating for failed codeclimate check.	8 years ago
Jeremy Meulemans	f70b52d0b6	Adding gauge for number of open head chunks. Fixes #1710	8 years ago
Julius Volz	beb3c4b389	Remove legacy remote storage implementations This removes legacy support for specific remote storage systems in favor of only offering the generic remote write protocol. An example bridge application that translates from the generic protocol to each of those legacy backends is still provided at: documentation/examples/remote_storage/remote_storage_bridge See also https://github.com/prometheus/prometheus/issues/10 The next step in the plan is to re-add support for multiple remote storages.	8 years ago
beorn7	d771185a43	storage: Fix chunkIndexToStartSeek calculation With a high enough shrink ratio and enough chunks to persist, the cutoff point could be _outside_ of the file, which wreaks havoc in the storage.	8 years ago
beorn7	73bd5e4dff	Merge branch 'beorn7/storage' into beorn7/storage3	8 years ago
beorn7	46a0837816	storage: Fix offset returned by dropAndPersistChunks This is another corner-case that was previously never exercised because the rewriting of a series file was never prevented by the shrink ratio. Scenario: There is an existing series on disk, which is archived. If a new sample comes in for that file, a new chunk in memory is created, and the chunkDescsOffset is set to -1. If series maintenance happens before the series has at least one chunk to persist _and_ an insufficient chunks on disk is old enough for purging (so that the shrink ratio kicks in), dropAndPersistChunks would return 0, but it should return the chunk length of the series file.	8 years ago
beorn7	9d12204da5	Merge branch 'release-1.5'	8 years ago
beorn7	bed4934224	storage: One more persist error code path discovered Also, in that code path, set chunkDescsOffset to 0 rather than -1 in case of "dropped more chunks from persistence than from memory" so that no other weird things happen before the series is quarantined for good.	8 years ago
beorn7	242d8edcb5	Merge branch 'release-1.5'	8 years ago
beorn7	8c8baaa558	storage: writeMemorySeries needs to return true for quarantined series This is another fallout of my bug hunt.	8 years ago
Mitsuhiro Tanda	be8b1eb656	storage: optimize dropping chunks by using minShrinkRatio (#2397 ) storage: prevent unnecessary chunk header reading if minShrinkRatio > 0	8 years ago
beorn7	2363a90adc	storage: Do not throw away fully persisted memory series in checkpointing	8 years ago
Fabian Reinartz	ea3ba338dd	main: add flags for new storage	8 years ago
beorn7	244a65fb29	storage: Increase persist watermark before calling append The append call may reuse cds, and thus change its len. (In practice, this wouldn't happen as cds should have len==cap. Still, the previous order of lines was problematic.)	8 years ago
beorn7	75282b27ba	storage: Added checks for invariants	8 years ago
beorn7	31e9db7f0c	storage: Simplify evictChunkDesc method	8 years ago
Fabian Reinartz	5772f1a7ba	retrieval/storage: adapt to new interface This simplifies the interface to two add methods for appends with labels or faster reference numbers.	8 years ago
beorn7	65dc8f44d3	storage: Test for errors returned by MaybePopulateLastTime	8 years ago
beorn7	752fac60ae	storage: Remove race condition from TestLoop	8 years ago
beorn7	4ccfc93dcf	storage: Set shrink ratio in the constructor.	8 years ago
beorn7	b2f086c6c4	storage: Expose bug of not setting the shrink ratio in the contstructor	8 years ago
Brian Brazil	c1b547a90e	Only checkpoint chunkdescs and series that need persisting. (#2340 ) This decreases checkpoint size by not checkpointing things that don't actually need checkpointing. This is fully compatible with the v2 checkpoint format, as it makes series appear as though the only chunksdescs in memory are those that need persisting.	8 years ago
Fabian Reinartz	c691895a0f	retrieval: cache series references, use pkg/textparse With this change the scraping caches series references and only allocates label sets if it has to retrieve a new reference. pkg/textparse is used to do the conditional parsing and reduce allocations from 900B/sample to 0 in the general case.	8 years ago
Brian Brazil	f64c231dad	Allow checkpoints and maintenance to happen concurrently. (#2321 ) This is essential on larger Prometheus servers, as otherwise checkpoints prevent sufficient persisting of chunks to disk.	8 years ago
Fabian Reinartz	ad9bc62e4c	storage: extend appender and adapt it	8 years ago
Brian Brazil	1dcb7637f5	Add various persistence related metrics (#2333 ) Add metrics around checkpointing and persistence * Add a metric to say if checkpointing is happening, and another to track total checkpoint time and count. This breaks the existing prometheus_local_storage_checkpoint_duration_seconds by renaming it to prometheus_local_storage_checkpoint_last_duration_seconds as the former name is more appropriate for a summary. * Add metric for last checkpoint size. * Add metric for series/chunks processed by checkpoints. For long checkpoints it'd be useful to see how they're progressing. * Add metric for dirty series * Add metric for number of chunks persisted per series. You can get the number of chunks from chunk_ops, but not the matching number of series. This helps determine the size of the writes being made. * Add metric for chunks queued for persistence Chunks created includes both chunks that'll need persistence and chunks read in for queries. This only includes chunks created for persistence. * Code review comments on new persistence metrics.	8 years ago
Fabian Reinartz	304cae9928	tsdb: Use PartitionedDB constructor	8 years ago
Brian Brazil	f9e581907a	Make index queue bigger. (#2322 ) When a large Prometheus starts up fresh it can take many minutes to warmup and clear out the index queue. A larger queue means less blocking, bigger batches and cuts down startup time by ~50%.	8 years ago
Fabian Reinartz	bc20d93f0a	storage: rename iterator value getters to At()	8 years ago
Fabian Reinartz	7322c46b8e	storage: add mock iterator for test	8 years ago
Fabian Reinartz	f8fc1f5bb2	*: migrate ingestion to new batch Appender	8 years ago
Fabian Reinartz	71fe0c58a8	promql: misc fixes	8 years ago
Mitsuhiro Tanda	7e369b9318	expose max memory chunks metrics (#2303 ) * expose max memory chunks metrics	8 years ago

1 2 3 4 5 ...

964 Commits (f04b1b5559a80a4fd1745cf891ce392a056460c9)