open-nomad

Author	SHA1	Message	Date
Danielle Tomlinson	4b4b85e3f4	allocwatcher: Cleanup new migrator/watcher interface	2018-12-11 13:12:35 +01:00
Danielle Tomlinson	83720575de	client: Unify handling of previous and preempted allocs	2018-12-11 13:12:35 +01:00
Michael Schurter	8808ab9cea	Merge pull request #4953 from hashicorp/b-script-context-wrapper consul: add ScriptExecutor context wrapper	2018-12-10 17:22:53 -08:00
Michael Schurter	4c5f3ae82c	Merge pull request #4952 from hashicorp/b-script-context consul: fix script checks exiting after 1 run	2018-12-10 17:22:15 -08:00
Danielle Tomlinson	dff7093243	client: Wait for preempted allocs to terminate When starting an allocation that is preempting other allocs, we create a new group allocation watcher, and then wait for the allocations to terminate in the allocation PreRun hooks. If there's no preempted allocations, then we simply provide a NoopAllocWatcher.	2018-12-11 00:59:18 +01:00
Danielle Tomlinson	2cdef6a7b4	allocwatcher: Add Group AllocWatcher The Group Alloc watcher is an implementation of a PrevAllocWatcher that can wait for multiple previous allocs before terminating. This is to be used when running an allocation that is preempting upstream allocations, and thus only supports being ran with a local alloc watcher. It also currently requires all of its child watchers to correctly handle context cancellation. Should this be a problem, it should be fairly easy to implement a replacement using channels rather than a waitgroup. It obeys the PrevAllocWatcher interface for convenience, but it may be better to extract Migration capabilities into a seperate interface for greater clarity.	2018-12-11 00:58:27 +01:00
Alex Dadgar	457c6eb398	typo	2018-12-10 15:35:26 -08:00
Alex Dadgar	508a3dfa49	merge 087 and 090 changelog	2018-12-10 15:34:21 -08:00
Mahmood Ali	fa9b9028a5	Use max 3 precision in displaying floats When formating floats in `nomad node status`, use a maximum precision of 3.	2018-12-10 12:18:24 -05:00
Mahmood Ali	14668f48d1	device attributes in `nomad node status -verbose` This reports device attributes like the following: ``` $ nomad node status -self -verbose ID = f7adb958-29e1-2a5a-2303-9d61ffaab33a Name = mars.local Class = <none> DC = dc1 Drain = false Eligibility = eligible Status = ready Uptime = 12h40m13s Drivers Driver Detected Healthy Message Time docker true true healthy 2018-12-10T11:47:19-05:00 ... Attributes cpu.arch = amd64 cpu.frequency = 2200 cpu.modelname = Intel(R) Core(TM) i7-8750H CPU @ 2.20GHz cpu.numcores = 12 ... Device Group Attributes Device Group = nomad/file/mock block_device = sda1 filesystem = ext4 size = 63.2 GB Meta ```	2018-12-10 12:18:24 -05:00
Mahmood Ali	9f69b8bfec	Rename helper_stats -> helper_devices	2018-12-10 12:18:24 -05:00
Mahmood Ali	ced76978bb	devices/nvidia: memory state as the summary stat	2018-12-10 12:18:24 -05:00
Mahmood Ali	97829a3f02	fix dtestutil.NewDriverHarness ref	2018-12-08 09:58:23 -05:00
Mahmood Ali	021d3720b5	Merge pull request #4950 from hashicorp/b-exc-libcontainer-kill executor: kill all container processes	2018-12-08 09:52:42 -05:00
Nick Ethier	32057b6f7f	Merge pull request #4973 from emate/recover-filerotator-from-io-errors Recover from any possible io error when invoking Write on FileRotator	2018-12-08 00:05:42 -05:00
Preetha Appan	f60c52c8ba	Score combinations of allocs from multiple devices for preemption	2018-12-07 18:35:47 -06:00
Chris Baker	4bbb8106c1	updated memberlist dependency to latest, which is missing NMD-1173 error	2018-12-07 22:15:05 +00:00
Chris Baker	59beae35df	nomad/rpc listener: modified to throttle logging on "permanent" Accept() errors as well (with a higher delay cap)	2018-12-07 22:14:15 +00:00
Alex Dadgar	695fa416a6	Merge pull request #4965 from hashicorp/b-gc-running Don't GC running but desired stop allocations	2018-12-07 13:36:33 -08:00
Chris Baker	707bac0a7b	rpc accept loop: added backoff on logging for failed connections, in case there is a fast fail loop (NMD-1173)	2018-12-07 20:12:55 +00:00
Marcin Matlaszek	39eec70f31	Recover from any possible io error when invoking Write on FileRotator As of now, FileRotator uses bufio.Write under the hood to write data to configured output file. Due to the way how bufio handles any occurred io error - saves it into `err` variable never resetting it automatically - any operation like `Write`, `Flush` etc will become a no-op, returning the very same, saved error (eg. Out of disk space) even when the problem is fixed (eg. disk space is available again). That automatically means that FileRotator will stop writing any logs, reporting the same error over and over again, even if it's no longer valid. This PR fixes it by resetting the bufio Writer, which resets any errors and tries to write requested data.	2018-12-07 18:22:29 +01:00
Mahmood Ali	7d5b5bb5f9	Merge pull request #4933 from hashicorp/f-mount-device Mount Devices in container based drivers	2018-12-07 10:32:03 -05:00
Mahmood Ali	91a67f347d	Vendor libcontainer/devices	2018-12-07 09:13:27 -05:00
Alex Dadgar	c918a96490	Warn if IOPS is being used	2018-12-06 16:17:09 -08:00
Alex Dadgar	1e3c3cb287	Deprecate IOPS IOPS have been modelled as a resource since Nomad 0.1 but has never actually been detected and there is no plan in the short term to add detection. This is because IOPS is a bit simplistic of a unit to define the performance requirements from the underlying storage system. In its current state it adds unnecessary confusion and can be removed without impacting any users. This PR leaves IOPS defined at the jobspec parsing level and in the api/ resources since these are the two public uses of the field. These should be considered deprecated and only exist to allow users to stop using them during the Nomad 0.9.x release. In the future, there should be no expectation that the field will exist.	2018-12-06 15:09:26 -08:00
Danielle Tomlinson	8100252116	Merge pull request #4960 from hashicorp/dani/b-gc-tests Re-enable Client GC tests	2018-12-06 23:18:36 +01:00
Mahmood Ali	a7b205daf2	Merge pull request #4955 from hashicorp/fix-docker-tests-20181203 Fix docker driver tests	2018-12-06 16:41:33 -05:00
Danielle Tomlinson	e3621c55fa	gc: Fix maxallocs integration test	2018-12-06 21:50:50 +01:00
Mahmood Ali	9e825f880c	Use absolute path in example device plugin deviceDir is used for specifying mount/device host paths, and those should be absolute paths.	2018-12-06 15:46:35 -05:00
Mahmood Ali	bdc53b1d8e	driver/rkt: mount plugin devices	2018-12-06 15:46:35 -05:00
Mahmood Ali	2c0fd2a902	driver/lxc: mount plugin devices Also, LXC requires target paths to be relative. Container paths in LXC binds should never be absolute paths, so we strip any preceeding `/`, even if a user sets one.	2018-12-06 15:46:35 -05:00
Mahmood Ali	699875eb1c	fixup: add missed docker utils test	2018-12-06 15:46:35 -05:00
Mahmood Ali	e9557ae596	tests: ensure image is loaded as test setup	2018-12-06 15:36:43 -05:00
Alex Dadgar	c4b5f80918	Make alloc health watcher a postrun hook rather than shutdown hook	2018-12-06 12:30:31 -08:00
Michael Lange	81c2d8b4a2	Merge pull request #4967 from hashicorp/b-ui-stat-charts-can-escape-canvas UI: Keep line charts in their canvases at all times	2018-12-06 10:56:37 -08:00
Danielle Tomlinson	62b98e64ca	client/gc: Replace GC integration test with unit The previous integration test was broken during the client refactor, and it seems to be some sort of race with state updating. I'm going to try and construct a replacement test as part of work on performance, but for now, the underlying behaviour is still being tested.	2018-12-06 12:28:23 +01:00
Danielle Tomlinson	f6e474fd55	client: Re-enable GC tests	2018-12-06 12:28:23 +01:00
Danielle Tomlinson	d043532cb0	allocrunner: Basic test alloc runner	2018-12-06 12:28:23 +01:00
Michael Lange	795ea7eade	Grow the default 0 to 1 bounds to the domain of the data when necessary	2018-12-05 22:07:44 -08:00
Alex Dadgar	b18a0f77a2	Merge pull request #4966 from hashicorp/b-failure-event Fix various bugs with task events	2018-12-05 14:43:50 -08:00
Alex Dadgar	b39c21d49c	Fix various bugs with task events Fixes the following: * Emitting events when the task fails to start * Don't double emit events on task shutdown (nomad stop) * Don't emit a OOM kill metric unless actually OOM'd	2018-12-05 14:27:07 -08:00
Alex Dadgar	14a61ea3ea	Don't GC running but desired stop allocations This PR fixes an edge case where we could GC an allocation that was in a desired stop state but had not terminated yet. This can be hit if the client hasn't shutdown the allocation yet or if the allocation is still shutting down (long kill_timeout). Fixes https://github.com/hashicorp/nomad/issues/4940	2018-12-05 13:01:12 -08:00
Mahmood Ali	b55fb642f1	driver/docker: honor plugin devices	2018-12-04 21:31:28 -05:00
Mahmood Ali	a580cef986	refactor device manipulation	2018-12-04 20:55:59 -05:00
Mahmood Ali	3a18105d06	drivers/exec: refactor stop/kill tests Simplify the tests to do all assertions within the main goroutine and account for status propagation delay.	2018-12-04 20:34:43 -05:00
Mahmood Ali	adb4d69576	Merge pull request #4956 from hashicorp/b-vault-client-tweaks-followup server/vault: Lock Vault expiration tracking	2018-12-04 19:46:59 -05:00
Mahmood Ali	366f478f8f	Merge pull request #4959 from hashicorp/fix-rkt-tests-20181204 tests: fix rkt tests	2018-12-04 19:46:41 -05:00
Mahmood Ali	428d35a5a9	executor: Keep 0.8.6 exit code for wait() failures 0.8.6 uses exit code 1 when `proc.Wait()` fails: https://github.com/hashicorp/nomad/blob/v0.8.6/client/driver/executor/executor.go#L442	2018-12-04 19:38:25 -05:00
Mahmood Ali	8df9de6fd5	driver/rkt: use rkt environment The rkt command itself needs an environment with PATH set to find iptables.	2018-12-04 14:00:45 -05:00
Preetha	8068d9f64e	Merge pull request #4949 from hashicorp/b-neg-running-summary Add guards around subtracting summary count	2018-12-04 12:52:58 -06:00

1 2 3 4 5 ...

13518 commits