k3s/test at ba535d57f6e2eb338f6fbf121d6be2e6f9204136 - k3s

History

Kubernetes Submit Queue a3f40dd8df Merge pull request #60856 from jiayingz/race-fix Automatic merge from submit-queue. If you want to cherry-pick this change to another branch, please follow the instructions <a href="https://github.com/kubernetes/community/blob/master/contributors/devel/cherry-picks.md">here</a>. Fixes the races around devicemanager Allocate() and endpoint deletion. There is a race in predicateAdmitHandler Admit() that getNodeAnyWayFunc() could get Node with non-zero deviceplugin resource allocatable for a non-existing endpoint. That race can happen when a device plugin fails, but is more likely when kubelet restarts as with the current registration model, there is a time gap between kubelet restart and device plugin re-registration. During this time window, even though devicemanager could have removed the resource initially during GetCapacity() call, Kubelet may overwrite the device plugin resource capacity/allocatable with the old value when node update from the API server comes in later. This could cause a pod to be started without proper device runtime config set. To solve this problem, introduce endpointStopGracePeriod. When a device plugin fails, don't immediately remove the endpoint but set stopTime in its endpoint. During kubelet restart, create endpoints with stopTime set for any checkpointed registered resource. The endpoint is considered to be in stopGracePeriod if its stoptime is set. This allows us to track what resources should be handled by devicemanager during the time gap. When an endpoint's stopGracePeriod expires, we remove the endpoint and its resource. This allows the resource to be exported through other channels (e.g., by directly updating node status through API server) if there is such use case. Currently endpointStopGracePeriod is set as 5 minutes. Given that an endpoint is no longer immediately removed upon disconnection, mark all its devices unhealthy so that we can signal the resource allocatable change to the scheduler to avoid scheduling more pods to the node. When a device plugin endpoint is in stopGracePeriod, pods requesting the corresponding resource will fail admission handler. Tested: Ran GPUDevicePlugin e2e_node test 100 times and all passed now. What this PR does / why we need it: Which issue(s) this PR fixes (optional, in `fixes #<issue number>(, fixes #<issue_number>, ...)` format, will close the issue(s) when PR gets merged): Fixes https://github.com/kubernetes/kubernetes/issues/60176 Special notes for your reviewer: Release note: ```release-note Fixes the races around devicemanager Allocate() and endpoint deletion. ```		2018-03-12 02:50:13 -07:00
..
conformance	Adds daemonset conformance tests	2018-02-28 12:43:02 -08:00
e2e	Merge pull request #60997 from MrHohn/e2e-fix-cleanup-svc-regional	2018-03-12 00:03:38 -07:00
e2e_node	Merge pull request #60856 from jiayingz/race-fix	2018-03-12 02:50:13 -07:00
fixtures	Make Scale() for RC poll-based until #31345 is fixed	2018-02-27 13:10:38 +01:00
images	Prevent webhooks from affecting admission requests for webhooks	2018-03-05 16:35:52 -08:00
integration	Merge pull request #60557 from hanxiaoshuai/fixtodo0228	2018-03-01 06:03:40 -08:00
kubemark	Rollback etcd server version to 3.1.11 due to #60589	2018-03-08 13:07:15 +01:00
list	Autogenerated: hack/update-bazel.sh	2018-02-16 13:43:01 -08:00
soak	Autogenerated: hack/update-bazel.sh	2018-02-16 13:43:01 -08:00
typecheck	Add OWNERS file to test/typecheck/	2018-03-06 22:35:47 -08:00
utils	Merge pull request #59840 from jennybuckley/webhooks-on-webhooks	2018-03-05 19:09:33 -08:00
BUILD	Add test/typecheck, a fast typecheck for all build platforms.	2018-02-27 13:53:32 -08:00
OWNERS	Add cblecker to test/ approvers	2018-03-06 22:35:26 -08:00
test_owners.csv	fix all the typos across the project	2018-02-11 11:04:14 +08:00
test_owners.json	Remove all traces of federation	2017-10-26 13:37:37 -07:00