open-nomad

Commit Graph

Author	SHA1	Message	Date
Seth Hoenig	fd2a31a331	drivers/docker: detect arch for default infra_image The 'docker.config.infra_image' would default to an amd64 container. It is possible to reference the correct image for a platform using the `runtime.GOARCH` variable, eliminating the need to explicitly set the `infra_image` on non-amd64 platforms. Also upgrade to Google's pause container version 3.1 from 3.0, which includes some enhancements around process management. Fixes #8926	2020-09-23 13:54:30 -05:00
James Rasell	dab8282be5	Merge pull request #8589 from hashicorp/f-gh-5718 driver/docker: allow configurable pull context timeout setting.	2020-08-14 16:07:59 +02:00
James Rasell	bc42cd2e5e	driver/docker: allow configurable pull context timeout setting. Pulling large docker containers can take longer than the default context timeout. Without a way to change this it is very hard for users to utilise Nomad properly without hacky work arounds. This change adds an optional pull_timeout config parameter which gives operators the possibility to account for increase pull times where needed. The infra docker image also has the option to set a custom timeout to keep consistency.	2020-08-12 08:58:07 +01:00
Nick Ethier	e39574be59	docker: support group allocated ports and host_networks (#8623 ) * docker: support group allocated ports * docker: add new ports driver config to specify which group ports are mapped * docker: update port mapping docs	2020-08-11 18:30:22 -04:00
Mahmood Ali	5796719124	docker: disable host volume binding by default	2020-06-23 13:43:37 -04:00
Seth Hoenig	ad91ba865c	driver/docker: enable setting hard/soft memory limits Fixes #2093 Enable configuring `memory_hard_limit` in the docker config stanza for tasks. If set, this field will be passed to the container runtime as `--memory`, and the `memory` configuration from the task resource configuration will be passed as `--memory_reservation`, creating hard and soft memory limits for tasks using the docker task driver.	2020-06-01 09:22:45 -05:00
Mahmood Ali	d9543a1a80	tests: don't delete images after tests complete Fix some docker test flakiness where image cleanup process may contaminate other tests. A clean up process may attempt to delete an image while it's used by another test.	2020-05-26 18:53:24 -04:00
Mahmood Ali	2588b3bc98	cleanup driver eventor goroutines This fixes few cases where driver eventor goroutines are leaked during normal operations, but especially so in tests. This change makes few modifications: First, it switches drivers to use `Context`s to manage shutdown events. Previously, it relied on callers invoking `.Shutdown()` function that is specific to internal drivers only and require casting. Using `Contexts` provide a consistent idiomatic way to manage lifecycle for both internal and external drivers. Also, I discovered few places where we don't clean up a temporary driver instance in the plugin catalog code, where we dispense a driver to inspect and validate the schema config without properly cleaning it up.	2020-05-26 11:04:04 -04:00
Tim Gross	aa8927abb4	volumes: return better error messages for unsupported task drivers (#8030 ) When an allocation runs for a task driver that can't support volume mounts, the mounting will fail in a way that can be hard to understand. With host volumes this usually means failing silently, whereas with CSI the operator gets inscrutable internals exposed in the `nomad alloc status`. This changeset adds a MountConfig field to the task driver Capabilities response. We validate this when the `csi_hook` or `volume_hook` fires and return a user-friendly error. Note that we don't currently have a way to get driver capabilities up to the server, except through attributes. Validating this when the user initially submits the jobspec would be even better than what we're doing here (and could be useful for all our other capabilities), but that's out of scope for this changeset. Also note that the MountConfig enum starts with "supports all" in order to support community plugins in a backwards compatible way, rather than cutting them off from volume mounting unexpectedly.	2020-05-21 09:18:02 -04:00
Mahmood Ali	182b95f7b1	use allow_runtimes for consistency Other allow lists use allow_ prefix (e.g. allow_caps, allow_privileged).	2020-05-12 11:03:08 -04:00
Mahmood Ali	0d692f0931	Add a knob to restrict docker runtimes	2020-05-12 10:14:43 -04:00
Ben Buzbee	769a3cd8b3	Rename OCIRuntime to Runtime; allow gpu conflicts is they are the same runtime; add conflict test	2020-04-03 12:15:11 -07:00
Ben Buzbee	d4f26d1eee	Support custom docker runtimes This enables customers who want to use gvisor and have it configured on their clients.	2020-04-03 11:07:37 -07:00
John Schlederer	8b35c75206	Making pull activity timeout configurable in Docker * Making pull activity timeout configurable in Docker plugin config, first pass * Fixing broken function call * Fixing broken tests * Fixing linter suggestion * Adding documentation on new parameter in Docker plugin config * Adding unit test * Setting min value for pull_activity_timeout, making pull activity duration a private var	2019-12-18 12:58:53 +01:00
Mahmood Ali	46bc3b57e6	address review comments	2019-12-13 11:21:00 -05:00
Mahmood Ali	0b7085ba3a	driver: allow disabling log collection Operators commonly have docker logs aggregated using various tools and don't need nomad to manage their docker logs. Worse, Nomad uses a somewhat heavy docker api call to collect them and it seems to cause problems when a client runs hundreds of log collections. Here we add a knob to disable log aggregation completely for nomad. When log collection is disabled, we avoid running logmon and docker_logger for the docker tasks in this implementation. The downside here is once disabled, `nomad logs ...` commands and API no longer return logs and operators must corrolate alloc-ids with their aggregated log info. This is meant as a stop gap measure. Ideally, we'd follow up with at least two changes: First, we should optimize behavior when we can such that operators don't need to disable docker log collection. Potentially by reverting to using pre-0.9 syslog aggregation in linux environments, though with different trade-offs. Second, when/if logs are disabled, nomad logs endpoints should lookup docker logs api on demand. This ensures that the cost of log collection is paid sparingly.	2019-12-08 14:15:03 -05:00
Nick Ethier	729dd9018c	docker: set default cpu cfs period (#6737 ) * docker: set default cpu cfs period Co-Authored-By: Michael Schurter <mschurter@hashicorp.com>	2019-11-19 19:05:15 -05:00
Mahmood Ali	977b86f924	driver/docker: ensure that defaults are populated Looks like we may need to pass default literal at each layer to be able, so defaults are set properly.	2019-10-18 18:27:28 -04:00
Mahmood Ali	3aec7b56ea	Only start reconciler once in main driver driver.SetConfig is not appropriate for starting up reconciler goroutine. Some ephemeral driver instances are created for validating config and we ought not to side-effecting goroutines for those. We currently lack a lifecycle hook to inject these, so I picked the `Fingerprinter` function for now, and reconciler should only run after fingerprinter started. Use `sync.Once` to ensure that we only start reconciler loop once.	2019-10-18 14:43:23 -04:00
Mahmood Ali	8739cc2a62	refactor reconciler code and address comments	2019-10-17 09:42:23 -04:00
Mahmood Ali	aa59280edc	docker: periodically reconcile containers When running at scale, it's possible that Docker Engine starts containers successfully but gets wedged in a way where API call fails. The Docker Engine may remain unavailable for arbitrary long time. Here, we introduce a periodic reconcilation process that ensures that any container started by nomad is tracked, and killed if is running unexpectedly. Basically, the periodic job inspects any container that isn't tracked in its handlers. A creation grace period is used to prevent killing newly created containers that aren't registered yet. Also, we aim to avoid killing unrelated containters started by host or through raw_exec drivers. The logic is to pattern against containers environment variables and mounts to infer if they are an alloc docker container. Lastly, the periodic job can be disabled to avoid any interference if need be.	2019-10-17 08:36:01 -04:00
Danielle Lancashire	724586ba1d	docker: Fix driver spec hclspec.NewLiteral does not quote its values, which caused `3m` to be parsed as a nonsensical literal which broke the plugin loader during initialization. By quoting the value here, it starts correctly.	2019-09-03 08:53:37 +02:00
Zhiguang Wang	832df1091b	Add default value "3m" to image_delay, making it consistent with docs.	2019-09-02 16:40:00 +08:00
Nick Ethier	1dae42ab81	docker: allow configuration of infra image	2019-07-31 01:04:07 -04:00
Nick Ethier	1fc5f86a7c	docker: support shared network namespaces	2019-07-31 01:03:20 -04:00
Mahmood Ali	104869c0e1	drivers/docker: rename logging `type` to `driver` Docker uses the term logging `driver` in its public documentations: in `docker` daemon config[1], `docker run` arguments [2] and in docker compose file[3]. Interestingly, docker used `type` in its API [4] instead of everywhere else. It's unfortunate that Nomad used `type` modeling after the Docker API rather than the user facing documents. Nomad using `type` feels very non-user friendly as it's disconnected from how Docker markets the flag and shows internal representation instead. Here, we rectify the situation by introducing `driver` field and prefering it over `type` in logging. [1] https://docs.docker.com/config/containers/logging/configure/ [2] https://docs.docker.com/engine/reference/run/#logging-drivers---log-driver [3] https://docs.docker.com/compose/compose-file/#logging [4] https://docs.docker.com/engine/api/v1.39/#operation/ContainerCreate	2019-02-28 16:04:03 -05:00
Mahmood Ali	4def8529db	driver/docker: use BlockAttrs for storage_opts storage_opts is a new field in 0.9 cycle and doesn't have backward compatibility constraints.	2019-02-19 20:35:28 -05:00
Mahmood Ali	46cd3c3f55	drivers: restore port_map old json support This ensures that `port_map` along with other block like attribute declarations (e.g. ulimit, labels, etc) can handle various hcl and json syntax that was supported in 0.8. In 0.8.7, the following declarations are effectively equivalent: ``` // hcl block port_map { http = 80 https = 443 } // hcl assignment port_map = { http = 80 https = 443 } // json single element array of map (default in API response) {"port_map": [{"http": 80, "https": 443}]} // json array of individual maps (supported accidentally iiuc) {"port_map: [{"http": 80}, {"https": 443}]} ``` We achieve compatbility by using `NewAttr("...", "list(map(string))", false)` to be serialized to a `map[string]string` wrapper, instead of using `BlockAttrs` declaration. The wrapper merges the list of maps automatically, to ease driver development. This approach is closer to how v0.8.7 implemented the fields [1][2], and despite its verbosity, seems to perserve 0.8.7 behavior in hcl2. This is only required for built-in types that have backward compatibility constraints. External drivers should use `BlockAttrs` instead, as they see fit. [1] https://github.com/hashicorp/nomad/blob/v0.8.7/client/driver/docker.go#L216 [2] https://github.com/hashicorp/nomad/blob/v0.8.7/client/driver/docker.go#L698-L700	2019-02-16 11:37:33 -05:00
Mahmood Ali	f7102cd01d	tests: add hcl task driver config parsing tests (#5314 ) * drivers: add config parsing tests Add basic tests for parsing and encoding task config. * drivers/docker: fix some config declarations * refactor and document config parse helpers	2019-02-12 14:46:37 -05:00
Michael Schurter	3b84e08fa4	Merge pull request #5297 from hashicorp/b-docker-logging Docker: Fix logging config parsing	2019-02-11 06:57:52 -08:00
Gertjan Roggemans	94ca78354b	docker: Fix volume driver_config options spec (#5309 ) Fixes #5308	2019-02-11 09:18:44 -05:00
Michael Schurter	e1e4b10884	docker: fix logging config parsing Fixes https://groups.google.com/d/topic/nomad-tool/B3Uo6Kns2BI/discussion	2019-02-04 11:07:57 -08:00
Michael Schurter	32daa7b47b	goimports until make check is happy	2019-01-23 06:27:14 -08:00
Michael Schurter	be0bab7c3f	move pluginutils -> helper/pluginutils I wanted a different color bikeshed, so I get to paint it	2019-01-22 15:50:08 -08:00
Alex Dadgar	cdcd3c929c	loader and singleton	2019-01-22 15:11:57 -08:00
oleksii.shyman	e41fbf7577	Add support for docker runtimes - docker fingerprint issues a docker api system info call to get the list of supported OCI runtimes. - OCI runtimes are reported as comma separated list of names - docker driver is aware of GPU runtime presence - docker driver throws an error when user tries to run container with GPU, when GPU runtime is not present - docker GPU runtime name is configurable	2019-01-15 11:34:47 -08:00
Nick Ethier	7e306afde3	executor: fix failing stats related test	2019-01-12 12:18:23 -05:00
Mahmood Ali	9369b123de	use drivers.FSIsolation	2019-01-08 09:11:47 -05:00
Alex Dadgar	4c57d2ec4d	Add plugin API versioning to plugin loader and plugins	2018-12-18 16:48:00 -08:00
Mahmood Ali	990a7d6776	driver/docker: stopping a dead container not error	2018-12-15 15:03:56 -05:00
Mahmood Ali	a580cef986	refactor device manipulation	2018-12-04 20:55:59 -05:00
Mahmood Ali	f6d6a50c39	add support for tmpfs	2018-11-27 07:20:17 -05:00
Mahmood Ali	0a09f5521d	Support docker bind mounts	2018-11-27 07:20:17 -05:00
Mahmood Ali	e9e415f186	Add support for storage opt	2018-11-20 16:11:02 -05:00
Nick Ethier	3ccd359735	docker: unexport new coordinator func	2018-11-19 23:07:07 -05:00
Nick Ethier	8b9b2b476e	docker: add default blocks for driver plugin config schema	2018-11-19 22:59:18 -05:00
Nick Ethier	2667f48a5d	docker: move config RPCs to config.go	2018-11-19 22:59:18 -05:00
Nick Ethier	ce4b867d21	docker: move volume driver options to seperate block	2018-11-19 22:59:18 -05:00
Nick Ethier	fca2df3c79	docker: group common config into blocks	2018-11-19 22:59:17 -05:00
Nick Ethier	8ef73e63ce	docker: moved fingerprint code to it's own file	2018-11-19 22:59:17 -05:00

1 2

51 Commits