open-nomad

Author	SHA1	Message	Date
James Rasell	c071efbd6b	Merge pull request #11411 from hashicorp/f-gh-11406 cli: add json and template flag opts to acl bootstrap command.	2021-11-02 09:48:25 +01:00
James Rasell	afb6913428	rpc: set the deregistration eval priority to the job priority. Previously when creating an eval for job deregistration, the eval priority was set to the default value irregardless of the job priority. In situations where an operator would want to deregister a high priority job so they could re-register; the evaluation may get blocked for some time on a busy cluster because of the deregsiter priority. If a job had a lower than default priority and was deregistered, the deregister eval would get a priority higher than that of the job. If we attempted to register another job with a higher priority than this, but still below the default, the deregister would be actioned before the register. Both situations described above seem incorrect and unexpected from a user prespective. This fix modifies to behaviour to set the deregister eval priority to that of the job, if available. Otherwise the default value is still used.	2021-11-02 09:11:44 +01:00
James Rasell	9d0fe24e25	docs: document Consul timeout config parameter.	2021-11-02 08:28:45 +01:00
Charlie Voiselle	29e7d46dd9	Making RPC Upgrade mode reloadable. (#11144 ) - Making RPC Upgrade mode reloadable. - Add suggestions from code review - remove spurious comment - switch to require(t,...) form for test. - Add to changelog	2021-11-01 16:30:53 -04:00
Luiz Aoqui	655ac2719f	Allow using specific object ID on diff (#11400 )	2021-11-01 15:16:31 -04:00
Michael Schurter	efe5714840	core: bump rejected plans from debug -> info As we have continued to see reports of #9506 we need to elevate this log line as it is the only way to detect when plans are being erroneously rejected. Users who see this log line repeatedly should drain and restart the node in the log line. This seems to workaorund the issue. Please post any details on #9506!	2021-10-31 12:51:42 -07:00
James Rasell	e151d93c9c	Merge pull request #11408 from VaultVulp/patch-1 Fix typo in documentation	2021-10-29 09:44:58 +02:00
James Rasell	30ad7985b2	changelog: add entry for #11411 .	2021-10-29 09:08:10 +02:00
James Rasell	46564ac579	docs: update acl bootstrap command to show json and template opts.	2021-10-29 09:01:58 +02:00
James Rasell	6c9e6e6f20	cli: add json and template flag opts to acl boostrap command.	2021-10-29 09:00:50 +02:00
Pavel Alimpiev	068066cb0e	Fix typo in documentation	2021-10-29 03:31:53 +03:00
James Rasell	f880edf6cf	Merge pull request #11338 from hashicorp/f-expose-nomad-consul-vagrant-linux vagrantfile: expose Nomad and Consul APIs to local machine.	2021-10-28 14:13:16 +02:00
Dave May	509c74ce19	debug: update default node-id and docs (#11398 ) * debug: default node-id to all * debug: align cli help and website documentation	2021-10-27 13:43:56 -04:00
Mahmood Ali	cdddd64a42	logging: Log the cause behind agent startup failure (#11353 ) Log the failure error when the agent fails to start. Previously, the agent startup failure error would be emitted to the command UI but not logged. So it doesn't get emitted to syslog or `log_file` if they are set, and it makes debugging much harder. Also, logging the error again before exit makes the error more visible: previously, the operator needed to scroll to the top to find the error. On a sample failure, the output will look like: ``` ==> WARNING: Bootstrap mode enabled! Potentially unsafe operation. ==> Loaded configuration from sample-configs/config-bad ==> Starting Nomad agent... ==> Error starting agent: setting up server node ID failed: mkdir /path-without-permission: read-only file system 2021-10-20T14:38:51.179-0400 [WARN] agent.plugin_loader: skipping external plugins since plugin_dir doesn't exist: plugin_dir=/path-without-permission/plugins 2021-10-20T14:38:51.181-0400 [DEBUG] agent.plugin_loader.docker: using client connection initialized from environment: plugin_dir=/path-without-permission/plugins 2021-10-20T14:38:51.181-0400 [DEBUG] agent.plugin_loader.docker: using client connection initialized from environment: plugin_dir=/path-without-permission/plugins 2021-10-20T14:38:51.181-0400 [INFO] agent: detected plugin: name=java type=driver plugin_version=0.1.0 2021-10-20T14:38:51.181-0400 [INFO] agent: detected plugin: name=docker type=driver plugin_version=0.1.0 2021-10-20T14:38:51.181-0400 [INFO] agent: detected plugin: name=mock_driver type=driver plugin_version=0.1.0 2021-10-20T14:38:51.181-0400 [INFO] agent: detected plugin: name=raw_exec type=driver plugin_version=0.1.0 2021-10-20T14:38:51.181-0400 [INFO] agent: detected plugin: name=exec type=driver plugin_version=0.1.0 2021-10-20T14:38:51.181-0400 [INFO] agent: detected plugin: name=qemu type=driver plugin_version=0.1.0 2021-10-20T14:38:51.181-0400 [ERROR] agent: error starting agent: error="setting up server node ID failed: mkdir /path-without-permission: read-only file system" ``` This change adds the final `ERROR` message. It's easy to miss the `==> Error starting agent` above.	2021-10-27 10:41:17 -07:00
Mike Nomitch	569a55675b	Replaces accidental use of Vault with Nomad (#11355 )	2021-10-27 08:35:31 -07:00
Mahmood Ali	daf20f9788	vault: set JobID in Vault metadata (#11397 ) Closes: #11395 .	2021-10-27 07:20:29 -07:00
Mahmood Ali	e06ff1d613	scheduler: stop allocs in unrelated nodes (#11391 ) The system scheduler should leave allocs on draining nodes as-is, but stop node stop allocs on nodes that are no longer part of the job datacenters. Previously, the scheduler did not make the distinction and left system job allocs intact if they are already running. I've added a failing test first, which you can see in https://app.circleci.com/jobs/github/hashicorp/nomad/179661 . Fixes https://github.com/hashicorp/nomad/issues/11373	2021-10-27 07:04:13 -07:00
Mahmood Ali	f03d65062d	Fix arm64 panics by updating google/snappy library to latest, 0.0.4 (#11396 ) Pick up https://github.com/golang/snappy/pull/56 to handle arm64 architectures to fix panics. tldr; Golang 1.16 changed `memmove` implementation for arm64 requiring additional cpu registers that snappy wasn't preserving in its assembly implementation. Other projects have experienced this issue as well, searching for `encode_arm64.s:666` on your favorite search engine will reveal some. Vault updated the dependency earlier this August: https://github.com/hashicorp/vault/pull/12371 . I believe this issue affects Nomad 1.2.x and 1.1.x. Nomad 1.0.x use Golang 1.15 and isn't affected. However, backporting the change to 1.0.x should be harmless. Fixed https://github.com/hashicorp/nomad/issues/11385 .	2021-10-27 06:39:16 -07:00
James Rasell	e4f703b401	vagrantfile: expose Nomad and Consul APIs to local machine.	2021-10-27 12:15:37 +02:00
Luiz Aoqui	b463715a98	prevent active log from being overwritten when agent starts (#11386 )	2021-10-26 20:57:07 -04:00
Luiz Aoqui	ecc7a288ec	docs: add note and example of storing `nomad job plan` index to disk (#11377 )	2021-10-26 20:25:22 -04:00
Charlie Voiselle	7d02c8b605	DOCS: Update Consul Connect to Consul service mesh (#11362 ) * Update Consul Connect to Consul service mesh * Apply suggestions from code review	2021-10-26 15:10:21 -04:00
Noel Quiles	f16ef7f6fb	website: Add Fathom analytics (#11276 ) * Impl Fathom analytics * Actually install fathom-client * Use analytics package instead of direct impl * Remove explicit fathom-client dep * Upgrade platform analytics package	2021-10-25 15:23:38 -04:00
Luiz Aoqui	645a87f6b3	ui: update task group alloc summary chart to use new `SummaryLegendItem` component (#11375 )	2021-10-25 11:14:01 -04:00
Luiz Aoqui	979faf41e5	fix test names (#11374 )	2021-10-22 15:43:55 -04:00
Luiz Aoqui	3c22fc79a5	add dispatch idempotency token support in the CLI (#10930 )	2021-10-22 12:39:05 -04:00
Luiz Aoqui	2c7bfb7000	ui: persist node drain settings (#11368 )	2021-10-22 10:51:31 -04:00
Luiz Aoqui	dc5222f6e5	ui: display Nomad version in the Clients and Servers table (#11366 )	2021-10-22 10:33:06 -04:00
Luiz Aoqui	a7eb72f7d1	ui: use `get` to access job meta value (#11370 )	2021-10-22 10:05:48 -04:00
Luiz Aoqui	b73ecf684b	ui: update favicon (#11371 )	2021-10-22 09:40:38 -04:00
Luiz Aoqui	6853bf9632	cli: allow setting namespace and region in the `nomad ui` command (#11364 )	2021-10-21 16:24:39 -04:00
Luiz Aoqui	fce1a03897	ui: create tooltip component (#11363 )	2021-10-21 13:12:33 -04:00
Luiz Aoqui	362c8c54f4	ui: set * as the default namespace selector (#11357 )	2021-10-21 10:24:07 -04:00
Luiz Aoqui	dceeccfc5d	ui: add client name tooltip when displaying client ID in tables (#11358 )	2021-10-21 10:23:06 -04:00
James Rasell	6011411111	Merge pull request #11339 from hashicorp/b-website-fixup-interpolation-formatting website: fixup link formatting within interpolation doc.	2021-10-21 09:15:36 +02:00
Mahmood Ali	e992ebf58d	document GH-11346 fix (#11350 )	2021-10-20 22:03:19 -04:00
Brandon Romano	8c863288ed	Merge pull request #11356 from hashicorp/update-alert-banner Update HashiConf alert-banner expiration	2021-10-20 16:28:30 -07:00
Brandon Romano	5c4f4be3ca	Update HashiConf alert-banner expiration Updates the HashiConf Alert Banner expiration to 10/20 @ 11pm (PT)	2021-10-20 16:02:45 -07:00
Michael Schurter	37a8f27a35	Merge pull request #11331 from shishir-a412ed/init Add support for --init to docker driver.	2021-10-20 10:49:51 -07:00
Michael Schurter	f95f966e8b	Merge pull request #11347 from shishir-a412ed/cleanup Code cleanup: Remove extra if clause.	2021-10-20 09:37:10 -07:00
Mahmood Ali	1de395b42c	Fix preemption panic (#11346 ) Fix a bug where the scheduler may panic when preemption is enabled. The conditions are a bit complicated: A job with higher priority that schedule multiple allocations that preempt other multiple allocations on the same node, due to port/network/device assignments. The cause of the bug is incidental mutation of internal cached data. `RankedNode` computes and cache proposed allocations in https://github.com/hashicorp/nomad/blob/v1.1.6/scheduler/rank.go#L42-L53 . But scheduler then mutates the list to remove pre-emptable allocs in https://github.com/hashicorp/nomad/blob/v1.1.6/scheduler/rank.go#L293-L294, and `RemoveAllocs` mutates and sets the tail of cached slice with `nil`s triggering a nil-pointer derefencing case. I fixed the issue by avoiding the mutation in `RemoveAllocs` - the micro-optimization there doesn't seem necessary. Fixes https://github.com/hashicorp/nomad/issues/11342	2021-10-19 20:22:03 -04:00
Shishir Mahajan	dd93f72920	Code cleanup: Remove extra if clause. Signed-off-by: Shishir Mahajan <smahajan@roblox.com>	2021-10-19 16:52:11 -07:00
Michael Schurter	081cfb85d7	docs: add #11331 to changelog	2021-10-19 16:30:06 -07:00
Michael Schurter	fd68bbc342	test: update tests to properly use AllocDir Also use t.TempDir when possible.	2021-10-19 10:49:07 -07:00
Brandon Romano	4d3bdc0dbf	Merge pull request #11341 from hashicorp/nq.update-alert-banner-hcg2021-live website: Update alert banner for HashiConf	2021-10-19 07:01:04 -07:00
Michael Schurter	d25b60a82d	docs: add #11334 to changelog	2021-10-18 09:22:01 -07:00
Michael Schurter	10c3bad652	client: never embed alloc_dir in chroot Fixes #2522 Skip embedding client.alloc_dir when building chroot. If a user configures a Nomad client agent so that the chroot_env will embed the client.alloc_dir, Nomad will happily infinitely recurse while building the chroot until something horrible happens. The best case scenario is the filesystem's path length limit is hit. The worst case scenario is disk space is exhausted. A bad agent configuration will look something like this: ```hcl data_dir = "/tmp/nomad-badagent" client { enabled = true chroot_env { # Note that the source matches the data_dir "/tmp/nomad-badagent" = "/ohno" # ... } } ``` Note that `/ohno/client` (the state_dir) will still be created but not `/ohno/alloc` (the alloc_dir). While I cannot think of a good reason why someone would want to embed Nomad's client (and possibly server) directories in chroots, there should be no cause for harm. chroots are only built when Nomad runs as root, and Nomad disables running exec jobs as root by default. Therefore even if client state is copied into chroots, it will be inaccessible to tasks. Skipping the `data_dir` and `{client,server}.state_dir` is possible, but this PR attempts to implement the minimum viable solution to reduce risk of unintended side effects or bugs. When running tests as root in a vm without the fix, the following error occurs: ``` === RUN TestAllocDir_SkipAllocDir alloc_dir_test.go:520: Error Trace: alloc_dir_test.go:520 Error: Received unexpected error: Couldn't create destination file /tmp/TestAllocDir_SkipAllocDir1457747331/001/nomad/test/testtask/nomad/test/testtask/.../nomad/test/testtask/secrets/.nomad-mount: open /tmp/TestAllocDir_SkipAllocDir1457747331/001/nomad/test/.../testtask/secrets/.nomad-mount: file name too long Test: TestAllocDir_SkipAllocDir --- FAIL: TestAllocDir_SkipAllocDir (22.76s) ``` Also removed unused Copy methods on AllocDir and TaskDir structs. Thanks to @eveld for not letting me forget about this!	2021-10-18 09:22:01 -07:00
Noel Quiles	ef533b6e3b	Update alert banner for HashiConf Final cleanup/closer exp date	2021-10-18 11:52:29 -04:00
James Rasell	2f5f6e0fdd	website: fixup link formatting within interpolation doc.	2021-10-18 12:21:05 +02:00
Andy Assareh	8c638217ac	exactly one of ingress, terminating, or mesh must be configured i believe mesh should be included in this statement was omitted.	2021-10-15 14:15:02 -07:00

1 2 3 4 5 ...

21976 commits