node_exporter

Commit Graph

Author	SHA1	Message	Date
Johannes 'fish' Ziemke	6ea0aa73e4	Rename interface to device in netclass collector (#1224 ) * Rename interface to device in netclass collector This makes it consistent with other networking metrics like node_network_receive_bytes_total This closes #1223 Signed-off-by: Johannes 'fish' Ziemke <github@freigeist.org>	6 years ago
Ralf Horstmann	3867ad5ab0	Add diskstats collector for OpenBSD (#1250 ) * Add diskstats collector for OpenBSD Tested on i386 and amd64, OpenBSD 6.4 and -current. * Refactor diskstats collectors This moves common descriptors from Linux, Darwin, OpenBSD diskstats collectors into diskstats_common.go Signed-off-by: Ralf Horstmann <ralf+github@ackstorm.de>	6 years ago
David O'Rourke	d442108d7a	collector: Implement uname collector for FreeBSD (#1239 ) * collector: Implement uname collector for FreeBSD Signed-off-by: David O'Rourke <david.orourke@gmail.com>	6 years ago
Paul Gier	2b81bff518	collector: use path/filepath for handling file paths (#1245 ) Similar to #1228. Update the remaining collectors to use 'path/filepath' intead of 'path' for manipulating file paths. Signed-off-by: Paul Gier <pgier@redhat.com>	6 years ago
Ralf Horstmann	dda51ad06a	Fix staticcheck ST1003 warnings (#1249 ) This fixes a few staticcheck ST1003 warnings in OpenBSD CPU collector. No functional change. Signed-off-by: Ralf Horstmann <ralf+github@ackstorm.de>	6 years ago
mknapphrt	7fbdd0ae93	Update procfs vendor (#1248 ) Signed-off-by: Mark Knapp <mknapp@hudson-trading.com>	6 years ago
Paul Gier	40dce45d8d	collector/systemd: add new label "type" for systemd_unit_state (#1229 ) Adds a new label called "type" systemd_unit_state which contains the Type field from the unit file. This applies only to the .service and .mount unit types. The other unit types do not include the optional type field. Fixes #1210 Signed-off-by: Paul Gier <pgier@redhat.com>	6 years ago
Matt Layher	3b5c2f6463	collector: use path/filepath for handling file paths (#1228 ) Signed-off-by: Matt Layher <mdlayher@gmail.com>	6 years ago
Jon Davies	e766485286	Add kstat-based Solaris metrics (#1197 ) * collector/loadavg_solaris.go: Use libkstat to gather load averages. * go.mod: Added go-kstat. * boot_time_solaris.go: Added. * cpu_solaris.go: Added. * README.md: Updated entries for Solaris. * collector/zfs_solaris.go: Added. * CHANGELOG.md: Added note about kstat-based Solaris metrics. Signed-off-by: Jonathan Davies <jpds@protonmail.com>	6 years ago
Ben Kochie	070e4b2e17	Update Makefile.common (#1220 ) * Update Makefile.common Update to new staticcheck method[0]. [0]: https://github.com/prometheus/prometheus/pull/5057 Signed-off-by: Ben Kochie <superq@gmail.com> * Fix staticcheck errors. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Ben Kochie	73ddf5f1f7	netstat: Add TCP In/Out Segs (#1185 ) * netstat: Add TCP In/Out Segs In order to get a better idea of TCP packet loss, we need to know how many `node_netstat_Tcp_OutSegs` there are so we can compare this to `node_netstat_Tcp_RetransSegs`. Signed-off-by: Ben Kochie <superq@gmail.com> * Update fixtures Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Tariq Ibrahim	6bd51269b7	update to host_statistics64 for Darwin meminfo (#1183 ) Signed-off-by: tariqibrahim <tariq181290@gmail.com>	6 years ago
Ben Kochie	4abc6fba7d	Add fallback for missing /proc/1/mounts (#1172 ) * Add fallback for missing /proc/1/mounts On some systems, `/proc/1/mounts` is hidden from non-root users due to the `hidepid` procfs feature. Attempt to fallback to `/proc/mounts` if `/proc/1/mounts` is not found. Signed-off-by: Ben Kochie <superq@gmail.com> * Add tests. Signed-off-by: Ben Kochie <superq@gmail.com> * Add CHANGELOG entry. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Jerome Froelich	0cb0c4d911	Remove unused variable readOnly from filesystem_linux.go. (#1173 ) The pull request #1002 changed the logic used on Linux servers to determine if a filesystem is read-only. As a result of this change, the variable `readOnly` is now unused and can be removed. Signed-off-by: Jerome Froelich <jeromefroelich@hotmail.com>	6 years ago
Nemikolh	62f99f95f0	Add receive/transmit bytes total metric (wifi collector). (#1150 ) Signed-off-by: Nemikolh <Nemikolh@users.noreply.github.com>	6 years ago
ioriveur	17fee8081f	Check BSD's mib which accounts for swap size (#1149 ) * Change Dfly's CPU counting frequency, see: https://github.com/prometheus/node_exporter/issues/1129 Signed-off-by: iori-yja <fivio.11235813@gmail.com> * Convert Dfly's CPU unit into second Signed-off-by: iori-yja <fivio.11235813@gmail.com> * Check BSD's mib which accounts for swap size; see #1127 Signed-off-by: iori-yja <fivo.11235813@gmail.com> * fix swap check code Signed-off-by: iori-yja <fivo.11235813@gmail.com>	6 years ago
Arno Uhlig	6edd9d217e	[systemd] collect taskCurrent, tasksMax per systemd unit (#1098 ) * [systemd] collect taskCurrent, tasksMax per systemd unit Signed-off-by: Arno Uhlig <arno.uhlig@sap.com>	6 years ago
Ben Kochie	b1eec66640	Add TCPSynRetrans to netstat default filter (#1143 ) Tcp SYN packet retransmits are a very useful signal as they affect network performance disproportionately to regular TCP retransmits. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Matt Layher	073e056121	Merge pull request #1131 from prometheus/mdl-collector-export collector: export NodeCollector for documentation purposes	6 years ago
Matt Layher	c0a55e3f80	collector: add bounds check and test for filesystem collector (#1133 ) Signed-off-by: Matt Layher <mdlayher@gmail.com>	6 years ago
Patrick	bdc0e7e678	Collect additional common Infiniband counters (#1120 ) * Collect additional common Infiniband counters Signed-off-by: Patrick Freeman <will.pat.free@gmail.com>	6 years ago
Paul Gier	988f049040	collector/hwmon_linux: handle temperature sensor file which doesn't have item suffix (#1123 ) In some cases the file might be called "temp" instead of the usual format "temp<index>_<item>" as described in the kernel docs: https://www.kernel.org/doc/Documentation/hwmon/sysfs-interface In this case, treat this as an _input file containing the current temperature reading. Fixes #1122 Signed-off-by: Paul Gier <pgier@redhat.com>	6 years ago
Paul Gier	38163f234f	collector/diskstats: don't fail if there are extra stats, just ignore… (#1125 ) * collector/diskstats: don't fail if there are extra stats, just ignore them Signed-off-by: Paul Gier <pgier@redhat.com>	6 years ago
Matt Layher	778124a56c	collector: add bounds check and test for tcpstat collector (#1134 ) Signed-off-by: Matt Layher <mdlayher@gmail.com>	6 years ago
Matt Layher	3d798aa4a1	collector: fix golint problems in ZFS collector (#1132 ) Signed-off-by: Matt Layher <mdlayher@gmail.com>	6 years ago
Matt Layher	2c2ee93519	collector: export NodeCollector for documentation purposes Signed-off-by: Matt Layher <mdlayher@gmail.com>	6 years ago
Ben Kochie	a0a164defb	Update cpufreq metrics collector (#1117 ) * Update Linux cpufreq collector to use new procfs library functions. * Split thermal throttle collection to a separate function. * Add new required fixtures and repack ttar file. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Paul Gier	7057c64f45	fix a few minor golint warnings (#1110 ) Signed-off-by: Paul Gier <pgier@redhat.com>	6 years ago
Paul Gier	e8d8199072	Update diskstats for linux kernel 4.19 (#1109 ) The format of /proc/diskstats is changing in linux-4.19 to include some additional fields. See: https://www.kernel.org/doc/Documentation/iostats.txt * collector/diskstats: use constants for some hard coded strings * collector/diskstats: update diskstats for linux-4.19 * collector/diskstats: remove kernel doc url from individual metrics Signed-off-by: Paul Gier <pgier@redhat.com>	6 years ago
Ben Kochie	0880d460d7	Ignore additional virtual filesystems (#1104 ) Add more virtual filesystems to the default ignore list * bpf * cgroup2 * selinuxfs * squashfs Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Dario Maiocchi	01ec8c5c5c	Remove continue with label (#1084 ) Instead of continue with label use helper function Signed-off-by: dmaiocchi <dmaiocchi@suse.com>	6 years ago
Ben Kochie	a1ce712e22	Cleanup unused /proc/mounts fixture. (#1097 ) * Cleanup unused /proc/mounts fixture. * Ignore Uint -> Unit in codespell. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Mario Trangoni	3659260b66	infiniband: Handle iWARP* RDMA modules N/A (#974 ) * infiniband: Add not connected i40iw0/ports/1 fixtures * infiniband: Handle issue when iWARP* RDMA modules are not available This is related to #966, and handle this error, Jun 07 13:33:24 hostname node_exporter[81888]: time="2018-06-07T13:33:24+02:00" level=error msg="ERROR: infiniband collector failed after 0.000929s: strconv.ParseUint: parsing \"N/A (no PMA)\": invalid syntax" source="collector.go:132" Signed-off-by: Mario Trangoni <mjtrangoni@gmail.com>	6 years ago
Yecheng Fu	0f9842f20a	[continue 912] strip rootfs prefix for run in docker (#1058 ) * strip rootfs prefix for run in docker * Use `/` as default value of path.rootfs, and parse mounts from `/proc/1/mounts`. * No need to mount `/proc` and `/sys` because we share host's PID namespace, which allows processes within the container to see all of the processes on the system. Closes: #66 Signed-off-by: Ivan Mikheykin <ivan.mikheykin@flant.com> Signed-off-by: Yecheng Fu <cofyc.jackson@gmail.com>	6 years ago
Ralf Horstmann	9f820bd3ee	Update cpu collector for OpenBSD 6.4 (#1094 ) Starting with (not yet released) OpenBSD 6.4, sysctl KERN_CPTIME2 will return ENODEV for offline CPUs. SMT siblings are reported as offline when hw.smt is disabled, which is the default since one of the later Spectre variants. So this might affect a few systems. For more details see: https://cvsweb.openbsd.org/src/sys/kern/kern_sysctl.c#rev1.348 Signed-off-by: Ralf Horstmann <ralf+github@ackstorm.de>	6 years ago
Daniele Sluijters	d999dacdc6	filesystem: Ignore netns/nsfs mounts (#1047 ) When starting Docker containers a whole bunch of netns (network namespace) mounts are created that the node exporter can't make any sense of (and can't read either). This ignores all nsfs filesystems. Fixes #875 Signed-off-by: Daniele Sluijters <daenney@users.noreply.github.com>	6 years ago
Ben Kochie	0fdc089187	Change systemd unit filtering (#1083 ) * Change systemd unit filtering Get all units from systemd and filter in Go. * Improves compatibility with older versions of systemd. * Improve debugging by printing when units pass the filter. * Remove extraneous newlines from log messages. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Luca Bruno	4672ea1671	collector/timex: remove cgo dependency (#1079 ) This removes the cgo import from timex collector, as it was only used to define two constants. Those are part of the Linux kernel<->userspace interface, thus there is no need to depend on libc to source them: https://github.com/torvalds/linux/blob/v4.18/include/uapi/linux/timex.h Signed-off-by: Luca Bruno <luca.bruno@coreos.com>	6 years ago
Björn Rabenstein	1c9ea46cca	Update vendoring for client_golang and friends (#1076 ) Signed-off-by: beorn7 <beorn@soundcloud.com>	6 years ago
Ben Kochie	ebdd524123	Correctly cast Darwin memory info (#1060 ) * Correctly cast Darwin memory info * Cast stats to float64 before doing math on them to avoid integer wrapping. * Remove invalid `_total` suffix from gauge values. * Handle counters in `meminfo.go`. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Marco Tulio R Braga	05e55bddad	Fix typo on description of read_time_seconds_total (#1057 ) Fix typo on unit description of metric `*read_time_seconds_total` from milliseconds to seconds. Signed-off-by: Marco Tulio R Braga <marco.tulio@mtulio.eng.br>	6 years ago
Dan Fredell	c52e0d3353	Fix SmartOS build #1017 (#1018 ) Signed-off-by: Dan Fredell <Dan.Fredell@gmail.com>	6 years ago
James Hartig	60c827231a	NRestarts or NRefused aren't available on older systemd versions (#1039 ) * If NRestarts or NRefused are not available, don't ignore the unit itself * Don't report systemd metrics (NRestarts/NRefused) that are not available Signed-off-by: James Hartig <james@getadmiral.com>	6 years ago
Ben Kochie	fe5a117831	Handle vanishing PIDs (#1043 ) PIDs can vanish (exit) from /proc/ between gathering the list of PIDs and getting all of their stats. * Ignore file not found errors. * Explicitly count the PIDs we find. * Cleanup some error style issues. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Ben Kochie	0662673ad6	Disable wifi collector by default (#1037 ) * Disable wifi collector by default Disable the wifi collector by default due to suspected cashing issues and goroutine leaks. * https://github.com/prometheus/node_exporter/issues/870 * https://github.com/prometheus/node_exporter/issues/1008 Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Ben Kochie	5d23ad0ca7	Fix supervisord collector (#978 ) * Replace supervisord xmlrpc library * Use `github.com/mattn/go-xmlrpc` that doesn't leak goroutines. * Fix uptime metric * Use Prometheus best practices for uptime metric. * Use "start time" rather than "uptime". * Don't emit a start time if the process is down. * Add changelog entry. * Add example compatibility rules. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Julius Volz	2c52b8c761	systemd: Remove unneeded/unhandled error returns (#1035 ) Signed-off-by: Julius Volz <julius.volz@gmail.com>	6 years ago
Christian Hoffmann	6bdc5558ec	build: make staticcheck happy by using real regexp patterns #1025 (#1026 ) Signed-off-by: Christian Hoffmann <mail@hoffmann-christian.info>	6 years ago
Hannes Körber	14a4f0028e	Enable nfs protocol (#998 ) * vendor: Update prometheus/procfs Signed-off-by: Hannes Körber <hannes.koerber@haktec.de> * mountstats: Use new NFS protocol field In https://github.com/prometheus/procfs/pull/100, the NFSTransportStats struct was expanded by a field called protocol that specifies the NFS protocol in use, either "tcp" or "udp". This commit adds the protocol as a label to all NFS metrics exported via the mountstats collector. Signed-off-by: Hannes Körber <hannes.koerber@haktec.de> * Update fixtures for UDP mount Signed-off-by: Hannes Körber <hannes.koerber@haktec.de>	6 years ago
Johannes Wienke	5c780d132c	Exclude only subdirectories of /var/lib/docker (#1003 ) It is quite common to put /var/lib/docker itself on a separate partition and that should be monitored as well. Signed-off-by: Johannes Wienke <languitar@semipol.de>	6 years ago
Ben Kochie	23f95c8e04	Fix ntp collector thread safety (#1014 ) Make the ntp collector thread safe by wrapping a mutex lock around the leapMidnight variable. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
xginn8	140b8b85c3	Filter out uninstalled systemd units when collecting all units (#1011 ) fixes #567 Signed-off-by: Matthew McGinn <mamcgi@gmail.com>	6 years ago
Sven Lange	2ae8c1c7a7	Add systemd uptime metric collection (#952 ) * Add systemd uptime metric collection Signed-off-by: Sven Lange <tdl@hadiko.de>	6 years ago
neiledgar	7e4d9bd150	Update wifi stats to support multiple stations (#977 ) (#980 ) Signed-off-by: neiledgar <neil.edgar@btinternet.com>	6 years ago
xginn8	9b97f44a70	Add a counter for refused socket unit connections, available as of systemd 239 (#995 ) Signed-off-by: xginn8 <mamcgi@gmail.com>	6 years ago
Brandon Gilmore	76bbd8dd18	Use /proc/mounts instead of statfs(2) for ro state (#1002 ) While the statfs(2) approach is reliable for normally mounted filesystems, the flags returned can be inconsistent when filesystem has been remounted read-only after encountering an error. The returned flags do accurately represent the internal state of the filesystem, but they do not reflect whether the VFS layer will accept writes. Instead, it makes sense to parse the current VFS mount state from the options field in /proc/mounts since it takes precedence. Signed-off-by: Brandon Gilmore <bgilmore@valvesoftware.com>	6 years ago
Jan Klat	c4102f1175	Add sys/class/net parsing from procfs and expose its metrics (#851 ) * add sys/class/net parsing from procfs and expose its metrics Signed-off-by: Jan Klat <jenik@klatys.cz> * change code to use int pointers per procfs change, move netclass to separate collector, change metric naming Signed-off-by: Jan Klat <jenik@klatys.cz> * bump year in licence, remove redundant newline, correct fixtures Signed-off-by: Jan Klat <jenik@klatys.cz> * fix style Signed-off-by: Jan Klat <jenik@klatys.cz> * change carrier changes to counter type Signed-off-by: Jan Klat <jenik@klatys.cz> * fix e2e output Signed-off-by: Jan Klat <jenik@klatys.cz> * add fixtures Signed-off-by: Jan Klat <jenik@klatys.cz> * update vendor, use fixtures correctly Signed-off-by: Jan Klat <jenik@klatys.cz> * change fixtures (device in /sys/class/net should be symlinked) Signed-off-by: Jan Klat <jenik@klatys.cz> * correct fixtures for 64k page, updated readme Signed-off-by: Jan Klat <jenik@klatys.cz>	6 years ago
mknapphrt	09b4305090	Changed the way that stuck mounts are handled. If a mount fails to return, it will stop being queried until it returns. (#997 ) Fixed spelling mistakes. Update transport_generic.go Changed to a mutex approach instead of channels and added a timeout before declaring a mount stuck. Removed unnecessary lock channel and clarified some var names. Fixed style nits. Signed-off-by: Mark Knapp <mknapp@hudson-trading.com>	6 years ago
xginn8	ac5a981761	Adding socket stat collection for systemd socket units (#968 ) Signed-off-by: xginn8 <mamcgi@gmail.com>	6 years ago
xginn8	8af84a215d	Add support for NRestarts counter introduced in systemd 235 (#992 ) * Add support for NRestarts counter introduced in systemd 235 `.service` units increment this counter any time the Restart= condition is triggered. Signed-off-by: Matthew McGinn <mamcgi@gmail.com>	6 years ago
Ben Kochie	107e5dfecc	Fix mdadm collector issues (#985 ) * Send "Personality unknown" to debug, not info, remove unnecessary newline. * Add support for "linear" personality. * Always set number of active disks to 0 when a device is inactive. * Add total disks calculation to unknown personalites. Signed-off-by: Ben Kochie <superq@gmail.com>	6 years ago
Derek Marcotte	2678d68dcc	Fix for #945 , cpu temperature is signed. (#965 ) * Fix for #945, cpu temperature is signed. Added a type conversion to cpu temperature sysctl. Will still collect/report -1 when the value is -1, this is because it should be up to interpretation whether this is the correct value for the system or not. Some drivers will report -1 for cpu temperature. Other sensors will report "an input into the fan control algorithm", i.e. not the actual temperature, but how much fan it wants. Some people cool their machines with liquid nitrogen. Signed-off-by: Derek Marcotte <554b8425@razorfever.net>	7 years ago
Brad Beam	e3cf1d5187	Adding support for evaluating octal characters in mountpoint (#954 ) Signed-off-by: Brad Beam <brad.beam@b-rad.info>	7 years ago
Pavlo Kutishchev	456bf5094a	Add processes exporter (#950 ) * Add processes exporter Signed-off-by: Pavel Kutishchev <pavel.kutishchev@olx.com> Signed-off-by: Ben Kochie <superq@gmail.com>	7 years ago
Alexey Kopytov	dd98a09bb2	A couple of ARM64-related fixes (#934 ) * Do not rely on AArch64 CPUs to support 32-bit ARM for cross-testing. Signed-off-by: Alexey Kopytov <akopytov@gmail.com> * aarch64 like ppc64le reports 64k node_sockstat_TCP_mem_bytes due to 64k pages. Signed-off-by: Alexey Kopytov <akopytov@gmail.com>	7 years ago
Steve Kotsopoulos	84dc362b05	Align Darwin disk stat names with Linux (#930 ) Signed-off-by: Steve Kotsopoulos <sk@fywss.com>	7 years ago
Mario Trangoni	24a28fcc9e	Remove unused func, var, and const (#928 ) Signed-off-by: Mario Trangoni <mjtrangoni@gmail.com>	7 years ago
Mario Trangoni	c9f421d0dd	Fix some golint issues (#927 ) * collector/cpu_: rename nodeCpuSecondsDesc to nodeCPUSecondsDesc Signed-off-by: Mario Trangoni <mjtrangoni@gmail.com> collector/qdisc_linux.go: add NewQdiscStatCollector comment Signed-off-by: Mario Trangoni <mjtrangoni@gmail.com> * collector/cpu_linux.go: rename core_map to coreMap Signed-off-by: Mario Trangoni <mjtrangoni@gmail.com>	7 years ago
Ben Kochie	361b5bf85d	Merge pull request #852 from prometheus/remove-gmond Remove gmond collector	7 years ago
Ben Kochie	b10ca77680	Fix /proc/net/dev/ interface name handling * Allow any character (UTF-8) for Linux interface names. Signed-off-by: Ben Kochie <superq@gmail.com>	7 years ago
Ben Kochie	1ab4a460c7	Update ppc64le end-to-end fixture. Signed-off-by: Ben Kochie <superq@gmail.com>	7 years ago
Johannes 'fish' Ziemke	fd66a86a30	Remove gmond collector Signed-off-by: Johannes 'fish' Ziemke <github@freigeist.org>	7 years ago
Ben Kochie	0f5be132ac	Merge pull request #904 from prometheus/superq/if_alias Fix parsing of interface aliases in netdev linux	7 years ago
Ben Kochie	a528966dcd	Fix parsing of interface aliases in netdev linux Very old kernels expose interface aliases as `foo0:0`, adjust the line parsing to handle these names. Signed-off-by: Ben Kochie <superq@gmail.com>	7 years ago
Ben Kochie	f6008b242b	Merge pull request #901 from mischief/bsd_boottime collector: implement node_boot_time_seconds for OpenBSD/NetBSD/Darwin	7 years ago
Jürgen Hötzel	de0632c2e9	Fix memory corruption when number of filesystems > 16 (#900 ) Signed-off-by: Juergen Hoetzel <juergen@archlinux.org>	7 years ago
mischief	26a385d7ab	collector: implement node_boot_time_seconds for OpenBSD/NetBSD/Darwin Signed-off-by: mischief <mischief@offblast.org>	7 years ago
Ben Kochie	015b86670a	Update ppc64le e2e output. Signed-off-by: Ben Kochie <superq@gmail.com>	7 years ago
Ben Kochie	0507b0c9a2	Fix formatting. Signed-off-by: Ben Kochie <superq@gmail.com>	7 years ago
Dmitriy Lukyanchikov	eddd1b9357	Fix netdev collector for linux (#890 ) fix variable name, fix transmitHeader extracting modify fixtures to run tests with updated netdev_linux collector Signed-off-by: dmitriy-lukyanchikov <d.lukyanchikov@anchorfree.com>	7 years ago
Derek Marcotte	fe86e908da	Update ppc64 fixtures to unbreak end-to-end. `efc1fdb` added new labels. Signed-off-by: Derek Marcotte <554b8425@razorfever.net>	7 years ago
Karsten Weiss	7e392e6634	Fix spelling mistakes found by codespell Signed-off-by: Karsten Weiss <knweiss@gmail.com>	7 years ago
Karsten Weiss	efc1fdb6d0	cpu: Add a 2nd label 'package' to metric node_cpu_core_throttles_total (#871 ) * cpu: Add a 2nd label 'package' to metric node_cpu_core_throttles_total This commit fixes the node_cpu_core_throttles_total metrics on multi-socket systems as the core_ids are the same for each package. I.e. we need to count them seperately. Rename the node_package_throttles_total metric label `node` to `package`. Reorganize the sys.ttar archive and use the same symlinks as the Linux kernel. Also, the new fixtures now use a dual-socket dual-core cpu w/o HT/SMT (node0: cpu0+1, node1: cpu2+3) as well as processor-less (memory-only) NUMA node 'node2' (this is a very rare case). Signed-off-by: Karsten Weiss <knweiss@gmail.com> * cpu: Use the direct /sys path to the cpu files. Use the direct path /sys/devices/system/cpu/cpu[0-9]* (without symlinks) instead of /sys/bus/cpu/devices/cpu[0-9]. The latter path also does not exist e.g. on RHEL 6.9's kernel. Signed-off-by: Karsten Weiss <knweiss@gmail.com> cpu: Reverse core+package throttle processing order Signed-off-by: Karsten Weiss <knweiss@gmail.com> * cpu: Add documentation URLs Signed-off-by: Karsten Weiss <knweiss@gmail.com>	7 years ago
Brian Brazil	31ce32f1fe	Greatly trim what netstat collector exposes by default (#876 ) Netstat is 40% of the metrics on my laptop, many of which are highly detailed information about IP internals in the kernel. ~300 such metrics on every machine in your fleet is excessive, so focus on key metrics by default, overridable by the user. Fixes #515 Signed-off-by: Brian Brazil <brian.brazil@robustperception.io>	7 years ago
Ben Kochie	cf3edadcbb	Update fixtures * Add oom_kill to fixture. * Update e2e outputs. * Put regexp in order. Signed-off-by: Ben Kochie <superq@gmail.com>	7 years ago
Brian Brazil	499c342fed	Greatly reduce the metrics vmstat returns by default. Vmstat has over 100 fields, most of which are highly detailed debug information. Trim this down to only essential fields by default, configurable by flag. Signed-off-by: Brian Brazil <brian.brazil@robustperception.io>	7 years ago
Brian Brazil	c8c144587e	Enable bonding collector by default. (#872 ) Signed-off-by: Brian Brazil <brian.brazil@robustperception.io>	7 years ago
Ben Kochie	779090db7e	Update ppc64le fixture (#867 ) Update to match standard e2e output. Signed-off-by: Ben Kochie <superq@gmail.com>	7 years ago
Mario Trangoni	1f11a86d59	Fix nfs golint issues (#863 ) * procfs: update vendoring Signed-off-by: Mario Trangoni <mjtrangoni@gmail.com> * procfs: fix e2e tests after nfs changes Signed-off-by: Mario Trangoni <mjtrangoni@gmail.com>	7 years ago
Ben Kochie	7b720df1c5	Use lowercase cpu label name in interrupts (#849 ) To match other CPU related metric labels, use a lowercase named label.	7 years ago
Johannes 'fish' Ziemke	424ca8e322	Drop exec_ in boot_timestamp_seconds on *bsd (#839 ) This closes #827.	7 years ago
colmbuckley	098f975b48	Correct the ClocksPerSec scaling factor on Darwin (#846 ) * Update cpu_darwin.go Change the definition of ClocksPerSec to read from limits.h * Update cpu_darwin.go	7 years ago
Julius Volz	864a6ee935	Treat custom textfile metric timestamps as errors (#769 ) This is clearer behavior and users will notice and fix their textfiles faster than if we just output a warning.	7 years ago
Rene Treffer	c504c7e264	Only report core throttles per core, not per cpu (#836 ) * Only report core throttles per core, not per cpu * Add topology/core_id to the cpu sysfs fixtures * Add new cpu fixtures to ttar file * Merge core_id reading and thermal throttle accounting * Declare core_id	7 years ago
Ben Kochie	e0d54a509c	Cleanup NFS metrics (#834 ) * Cleanup NFS metrics * Update `nfs` metric names to match `nfsd`. * Remove uneeded `tcp` label from TCP connections metric. * Remove uneeded `v` on `nfsd` metrics. * Enable all `nfs` v4 client metrics. * Remove `nfs` metric name overrides. * Add ppc64le fixture. * Fix typo.	7 years ago
Ben Kochie	3f41a2fecb	Update ppc64le fixture (#832 ) Updates fixture for ppc64le arch to latest output.	7 years ago
Ben Kochie	d33a447047	Remove deprecated prometheus.InstrumentHandlerFunc (#831 ) Update Prometheus client golang use to use `promhttp.Handler()` instead of `prometheus.InstrumentHandlerFunc()`.	7 years ago
Richard Elling	d7348a5c78	updates for zfsonlinux 0.7.5 (#779 ) * updates for zfsonlinux 0.7.5 * add constants for KSTAT_DATA_* types * added e2e test for negative values represented by uint64 that can result from ZFS bugs	7 years ago
Ben Kochie	6468e7c80b	Enable NFS client metrics by default. (#828 ) Enable NFS client metrics by default now that it nolonger prints errors on scrape if there are no metrics to display. Also fixup the nfsd README to match the nfs entry.	7 years ago
Ralf Horstmann	8d9c7ca659	Use swpginuse instead of swpgonly in meminfo_openbsd (#813 ) All tools in OpenBSD base system use swpginuse instead of swpgonly for reporting swap usage (snmpd, swapctl, top, vmstat), so let memory collector use that as well for consistency.	7 years ago
Kasinath Kottukkal	f6965e1812	Add overlay to defIgnoredFSTypes (#824 ) * Add overlay to defIgnoredFSTypes To avoid statfs() errors if node_exporter is running as non privileged user. * Updated defIngoredFSTypes values in sorted order	7 years ago
Ben Kochie	01bd99fb1a	Refactor NFS client collector (#816 ) * Update vendor github.com/prometheus/procfs/... * Refactor NFS collector Use new procfs library to parse NFS client stats. * Ignore nfs proc file not existing. * Refactor with reflection to walk the structs.	7 years ago
Brian Brazil	52c031890e	Add _seconds suffix to node_time. (#823 )	7 years ago
Ben Kochie	05eabe60fb	Fix error output in nfsd collector. (#821 )	7 years ago
Ben Kochie	3de2542d21	Fix NFSd metric type (#819 ) RPC Count should be a counter, not a gauge.	7 years ago
Matt Layher	544488ddd6	Fix remaining metric naming issues (#799 )	7 years ago
Ben Kochie	6a041692ed	Add NFS Server metrics collector. (#803 ) * Add NFS Server metrics collector. * Add File Handles metrics. * Add nfsd IO stats. * Add metrics for NFSd threads. * Add metrics for NFSd read ahead cache. * Add NFSd network traffic counters. * Add RPC metrics. * Add V2 requests metrics. * Add NFSv3 metrics. * Add NFSv4 metrics. * Update reply cache comment. * Update help text.	7 years ago
Brian Brazil	1072f2868d	Fix log level regression in #533	7 years ago
Brian Brazil	7e41a2b279	Ignore /var/lib/docker by default. (#814 ) The node exporter runs unprivileged, so it cannot statfs any filesystems under this directory causing log spam. In addition there tends to be high churn in the filesystems here (as it's basically application monitoring) which can cause high cardinaltiy and in one case caused Prometheus's index symbol table to get very large. Accordingly this should be ignored to reduce log spam and avoid performance issues. The filesystems themselves can in principle be monitored via container oriented exporters, and the underlying filesystems will still be monitored.	7 years ago
Ralf Horstmann	29ac809e48	Use unified CPU metric description on OpenBSD (#810 )	7 years ago
Derek Marcotte	fde5d2c6c9	Remove unsafe typecasts from sysctl_bsd getStructTimeval. (#741 ) There is a simpler way.	7 years ago
Ben Kochie	14d60958d6	Unify CPU collector conventions (#806 ) * Unify CPU collector conventions Add a common CPU metric description. * All collectors use the same `nodeCpuSecondsDesc`. * All collectors drop the `cpu` prefix for `cpu` label values. * Fix subsystem string in cpu_freebsd. * Fix Linux CPU freq label names.	7 years ago
Ralf Horstmann	e3c76b1f0c	Add OpenBSD CPU collector (#805 )	7 years ago
Tom Wilkie	6833eec187	Fix tests.	7 years ago
Tom Wilkie	0316bacceb	Only use one dbus connection, required some refactoring.	7 years ago
Tom Wilkie	a7fd6b8743	Export systemd timer last trigger sec.	7 years ago
Ben Kochie	111e3af437	Remove obsolete megacli collector. (#798 ) This collector has been replaced by the textfile collector tool `storcli.py`.	7 years ago
Julius Volz	6cac74f0e0	Add unit suffix to textfile collector mtime metric (#796 )	7 years ago
Brian Brazil	a98067a294	Make metrics better follow guidelines (#787 ) * Improve stat linux metric names. cpu is no longer used. * node_cpu -> node_cpu_seconds_total for Linux * Improve filesystem metric names with units * Improve units and names of linux disk stats Remove sector metrics, the bytes metrics cover those already. * Infiniband counters should end in _total * Improve timex metric names, convert to more normal units. See `3c073991eb/kernel/time/ntp.c (L909)` for what stabil means, looks like a moving average of some form. * Update test fixture * For meminfo metrics that had "kB" units, add _bytes * Interrupts counter should have _total	7 years ago
Ben Kochie	b4d7ba119a	Add fixture for ppc64le (#785 ) * Add support for per-architecture fixtures. * Add output for ppc64le.	7 years ago
Nick Owens	0629a081db	multiply page size after float64 coercion to avoid signed integer overflow (#780 )	7 years ago
Franz Pletz	d432f9857e	Use uint64 in the ZFS collector (#714 ) ZFS metrics can also be unsigned 64-bit integers that won't fit in int64 and causes the whole collector to fail.	7 years ago
Derek Marcotte	477fe4665a	Move FreeBSD/DragonflyBSD out of meminfo add kvm. (#547 ) * Move FreeBSD/DragonflyBSD out of meminfo add kvm. This gives us SwapUsed, and everything under one roof. * Fix typos per review. * Update to use newer API. * Remove premature optimization per PR feedback.	7 years ago
Sevag Hanssian	4329b0a86b	Add summary metrics for systemd exporter (#765 )	7 years ago
Matthieu Guegan	d6ef10bb56	Add openbsd meminfo (#724 ) * Implements meminfo collector for OpenBSD This is a rework of #151. * Fix CGO import * Add some useful metrics * Rename total -> size for normalization	7 years ago
Ben Kochie	7f6c59e198	Ignore more virtual filesystems (#775 ) Add additional Linux virtual filesystem types to the default list.	7 years ago
Netmonk	2aa8d0eb0c	[FIX] Exclude Linux proc from filesystem type regexp (#774 ) * [FIX] Issue 63, error on excluding proc filesystem on linux, improving regexp * [FIX] Reordering filter order	7 years ago
Julius Volz	f536857ac6	Fix e2e tests after textfile custom timestamp removal (#768 )	7 years ago
Shubheksha Jalan	1f2458f42c	Filter out testfile metrics correctly when using `collect[]` filters (#763 ) * remove injection hook for textfile metrics, convert them to prometheus format * add support for summaries * add support for histograms * add logic for handling inconsistent labels within a metric family for counter, gauge, untyped * change logic for parsing the metrics textfile * fix logic to adding missing labels * Export time and error metrics for textfiles * Add tests for new textfile collector, fix found bugs * refactor Update() to split into smaller functions * remove parseTextFiles(), fix import issue * add mtime metric directly to channel, fix handling of mtime during testing * rename variables related to labels * refactor: add default case, remove if guard for metrics, remove extra loop and slice * refactor: remove extra loop iterating over metric families * test: add test case for different metric type, fix found bug * test: add test for metrics with inconsistent labels * test: add test for histogram * test: add test for histogram with extra dimension * test: add test for summary * test: add test for summary with extra dimension * remove unnecessary creation of protobuf * nit: remove extra blank line	7 years ago
Ben Kochie	cd2a17176a	Add full make to CircleCI (#761 ) * Add full make to CircleCI Ensure end-to-end test is run. * Fix go fmt error. * Fix end-to-end output.	7 years ago
Wei Li	1e9bb4ec3a	textfile: fix duplicate metrics error (#738 ) The textfile gatherer should only be added to gatherer list once. Signed-off-by: Li Wei <liwei@anbutu.com>	7 years ago
Kristian Klausen	a96f1738b3	netdev: Change valueType to CounterValue (#749 ) All the metric only goes up, so the type should be counter. This also add _total to all the metric name. Fix: #747	7 years ago
Ben Kochie	2a80537547	Split out guest cpu metrics on Linux. (#744 ) Linux "guest" metrics for VMs are already accounted for in node_cpu `user` and `nice` metrics. Separate these into their own metric to avoid duplication of data.	7 years ago
Karsten Weiss	a8d7d1101a	cpu: Support processor-less (memory-only) NUMA nodes (#734 ) * cpu: Support processor-less (memory-only) NUMA nodes Processor-less (memory-only) NUMA nodes exist e.g. in systems that use Intel Optane drives for RAM expansion using Intel Memory Drive Technology (IMDT). IMDT RAM expansion supports two modes: * "Unify Remote Memory domains": present a processor-less (memory-only) NUMA domain, which is the default * "Expand local memory domains": to expand each processor’s memory domain with a portion of the memory made available by Optane and IMDT This commit fixes a crash in the first case (when "cpulist" is empty). Here's an example of such a system: $ numastat -m\|head -n5 Per-node system memory usage (in MBs): Node 0 Node 1 Node 2 Total --------------- --------------- --------------- --------------- MemTotal 118239.56 130816.00 464384.00 713439.56 $ for i in {0..2}; do echo -n "$i: " ; cat /sys/bus/node/devices/node$i/cpulist ; done 0: 0-7,16-23 1: 8-15,24-31 2: $ /opt/vsmp/bin/vsmpversion -vvv Memory Drive Technology: 8.2.1455.74 (Sep 28 2017 13:09:59) System configuration: Boards: 3 1 x Proc. + I/O + Memory 2 x NVM devices (Intel SSDPED1K375GAQ) Processors: 2, Cores: 16, Threads: 32 Intel(R) Xeon(R) CPU E5-2667 v4 @ 3.20GHz Stepping 01 Memory (MB): 713472 (of 977450), Cache: 251416, Private: 12562 1 x 249088MB [262036/ 678/12270] 1 x 232192MB [357707/125369/ 146] 82:00.0#1 1 x 232192MB [357707/125369/ 146] 83:00.0#1 * cpu: rename some variables (pkg => node) * cpu: Use %v not %q in log.Debugf() format strings	7 years ago
Matt Layher	f6f9c8d6cc	Add and use sysReadFile in hwmon collector (#728 )	7 years ago
Tobias Klauser	d73f1e60c4	Simplify Utsname string conversion (#716 ) * Update golang.org/x/sys/unix This allows to use simplified string conversion of Utsname members. * Simplify Utsname string conversion Use Utsname from golang.org/x/sys/unix which contains byte array instead of int8/uint8 array members. This allows to simplify the string conversions of these members.	7 years ago
Ben Kochie	ea250d73f4	Fix off by one in Linux interrupts collector (#721 ) * Fix off by one in Linux interrupts collector * Fix off by one in CPU column handler. * Add test. * Enable interrupts in end-to-end test.	7 years ago
Matt Layher	296b62acb7	netstat: return nothing when /proc/net/snmp6 not found	7 years ago
Derek Marcotte	0eecaa9547	Correct buffer_bytes > INT_MAX on BSD/amd64. (#712 ) * Correct buffer_bytes > INT_MAX on BSD/amd64. The sysctl vfs.bufspace returns either an int or a long, depending on the value. Large values of vfs.bufspace will result in error messages like: couldn't get meminfo: cannot allocate memory This will detect the returned data type, and cast appropriately. * Added explicit length checks per feedback. * Flatten Value() to make it easier to read. * Simplify per feedback. * Fix style. * Doc updates.	7 years ago
Matt Layher	f9ad88fc03	xfs: expose correct fields, fix metric names	7 years ago
Siavash Safi	f3a7022602	Add `collect[]` parameter (#699 ) * Add `collect[]` parameter * Add TODo comment about staticcheck ignored * Restore promhttp.HandlerOpts * Log a warning and return HTTP error instead of failing * Check collector existence and status, cleanups * Fix warnings and error messages * Don't panic, return error if collector registration failed * Update README	7 years ago
Ben Kochie	deadfef4c9	Update vendoring (#685 ) * Update vendor github.com/coreos/go-systemd/dbus@v15 * Update vendor github.com/ema/qdisc * Update vendor github.com/godbus/dbus * Update vendor github.com/golang/protobuf/proto * Update vendor github.com/lufia/iostat * Update vendor github.com/matttproud/golang_protobuf_extensions/pbutil@v1.0.0 * Update vendor github.com/prometheus/client_golang/... * Update vendor github.com/prometheus/common/... * Update vendor github.com/prometheus/procfs/... * Update vendor github.com/sirupsen/logrus@v1.0.3 Adds vendor golang.org/x/crypto * Update vendor golang.org/x/net/... * Update vendor golang.org/x/sys/... * Update end to end output.	7 years ago
Brett Vickers	b62c7bc0ad	Updated vendored ntp package (#681 ) The github.com/beevik/ntp package was recently updated with some API changes that broke node_exporter. This commit fetches the latest version of the ntp package and brings node_exporter in line with the latest API.	7 years ago
Calle Pettersson	859a825bb8	Replace --collectors.enabled with per-collector flags (#640 ) * Move NodeCollector into package collector * Refactor collector enabling * Update README with new collector enabled flags * Fix out-of-date inline flag reference syntax * Use new flags in end-to-end tests * Add flag to disable all default collectors * Track if a flag has been set explicitly * Add --collectors.disable-defaults to README * Revert disable-defaults flag * Shorten flags * Fixup timex collector registration * Fix end-to-end tests * Change procfs and sysfs path flags * Fix review comments	7 years ago
Sami Kerola	3762191e66	Add timex collector (#664 ) This collector is based on adjtimex(2) system call. The collector returns three values, status if time is synchronised, offset to remote reference, and local clock frequency adjustment. Values are taken from kernel time keeping data structures to avoid getting involved how the synchronisation is implemented. By that I mean one should not care if time is update using ntpd, systemd.timesyncd, ptpd, and so on. Since all time sync implementation will always end up telling to kernel what is the status with time one can simply omit the software in between, and look results of the syncing. As a positive side effect this makes collector very quick and conceptually specific, this does not monitor availability of NTP server, or network in between, or dns resolution, and other unrelated but necessary things. Minimum set of values to keep eye on are the following three: The node_timex_sync_status tells if local clock is in sync with a remote clock. Value is set to zero when synchronisation to a reliable server is lost, or a time sync software is misconfigured. The node_timex_offset_seconds tells how much local clock is off when compared to reference. In case of multiple time references this value is outcome of RFC 5905 adjustment algorithm. Ideally offset should be close to zero, and it depends about use case how large value is acceptable. For example a typical web server is probably fine if offset is about 0.1 or less, but that would not be good enough for mobile phone base station operator. The node_timex_freq tells amount of adjustment to local clock tick frequency. For example if offset is one second and growing the local clock will need instruction to tick quicker. Number value itself is not very important, and occasional small adjustments are fine. When frequency is unusually in stable one can assume quality of time stamps will not be accurate to very far in sub second range. Obviously explaining why local clock frequency behaves like a passenger in roller coaster is different matter. Explanations can vary from system load, to environmental issues such as a machine being physically too hot. Rest of the measurements can help when debugging. If you run a clock server do probably want to collect and keep track of everything. Pull-request: https://github.com/prometheus/node_exporter/pull/664	7 years ago
Leonid Evdokimov	c169b4b1c5	Add metrics from SNTPv4 packet to ntp collector & add ntpd sanity check (#655 ) * Add metrics from SNTPv4 packet to ntp collector & add ntpd sanity check 1. Checking local clock against remote NTP daemon is bad idea, local ntpd acting as a client should do it better and avoid excessive load on remote NTP server so the collector is refactored to query local NTP server. 2. Checking local clock against remote one does not check local ntpd itself. Local ntpd may be down or out of sync due to network issues, but clock will be OK. 3. Checking NTP server using sanity of it's response is tricky and depends on ntpd implementation, that's why common `node_ntp_sanity` variable is exported. * `govendor add golang.org/x/net/ipv4`, it is dependency of github.com/beevik/ntp * Update github.com/beevik/ntp to include boring SNTP fix * Use variable name from RFC5905 * ntp: move code to make export of raw metrics more explicit * Move NTP math to `github.com/beevik/ntp` * Make `golint` happy * Add some brief docs explaining `ntp` #655 and `timex` #664 modules * ntp: drop XXX comment that got its decision * ntp: add `_seconds` suffix to relevant metrics * Better `node_ntp_leap` comment * s/node_ntp_reftime/node_ntp_reference_timestamp_seconds/ as requested by @discordianfish * Extract subsystem name to const as suggested by @SuperQ	7 years ago
Karsten Weiss	b0d5c00832	cpu: Metric 'package_throttles_total' is per package. (#657 ) * cpu: Metric 'package_throttles_total' is per package. 'package_throttles_total' is per package, not per cpu. This also reduces the total number of cpu time series a lot (esp for multi core cpus). * cpu: Better handling of a cpulist edge-case. * cpu: Extract the package number from the directory name. Do not rely on the range index. * cpu: Add package_throttle_count for node0 cpu1 This file must be ignored by the cpu collector.	7 years ago
Matthias Rampke	e1f129c729	Use int64 throughout the ZFS collector. This avoids issues with integer overflows on 32-bit architectures. The Prometheus data format is float64, so regardless of the architecture we should handle large numbers. Fixes #629.	7 years ago
Ben Kochie	8839640cd1	Ignore wifi collector permission errors (#646 ) Ignore the permission denined error when the wifi collector has no permission to read metrics.	7 years ago
Calle Pettersson	dfe07eaae8	Switch to kingpin flags (#639 ) * Switch to kingpin flags * Fix logrus vendoring * Fix flags in main tests * Fix vendoring versions	7 years ago

1 2 3 4 5 ...

670 Commits (baa7ab732f34409d617fdab0b8eefa5e8dc4eaf6)