TPA, Ansible, and sudo
======================

TPA uses Ansible with sudo to execute tasks with elevated privileges on
target instances. This page explains how Ansible uses sudo (which is in
no way TPA-specific), and the consequences to systems managed with TPA.

TPA needs root privileges;

- to install packages (required packages using the operating system’s
  native package manager, and optional packages using pip)

- to stop, reload and restart services (i.e Postgres, repmgr, efm, etcd,
  haproxy, pgbouncer etc.)

- to perform a variety of other tasks (e.g., gathering cluster facts,
  performing switchover, setting up cluster nodes)

TPA also needs to be able to use sudo. You can make it ssh in as root
directly by setting ``ansible_user: root`` , but it will still use sudo
to execute tasks as other users (e.g., postgres).

Ansible sudo invocations
------------------------

When Ansible runs a task using sudo, you will see a process on the
target instance that looks something like this:

.. code:: shell

   /bin/bash -c sudo -H -S -n  -u root /bin/bash -c \
     ""echo BECOME-SUCCESS-kfoodiiprztsyerriqbjuqhhbemejgpc ; \
     /usr/bin/python2"" && sleep 0

People who were expecting something like ``sudo yum install -y xyzpkg``
are often surprised by this. By and large, most tasks in Ansible will
invoke a Python interpreter to execute Python code, rather than
executing recognisable shell commands. (Playbooks may execute ``raw``
shell commands, but TPA uses such tasks only to bootstrap a Python
interpreter.)

Ansible modules contain Python code of varying complexity, and an
Ansible playbook is not just a shell script written in YAML format.
There is no way to “extract” shell commands that would do the same thing
as executing an arbitrary Ansible playbook.

There is one significant consequence of how Ansible uses sudo:
`privilegeescalation must be general <https://docs.ansible.com/ansible/latest/playbook_guide/playbooks_privilege_escalation.html#privilege-escalation-must-be-general>`_  . That, it is not possible to limit sudo invocations to
specific commands in sudoers.conf, as some administrators are used to
doing. Most tasks will just invoke python. You could have restricted
sudo access to python if it were not for the random string in every
command—but once Python is running as root, there’s no effective limit
on what it can do anyway.

Executing Python modules on target hosts is just the way Ansible works.
None of this is specific to TPA in any way, and these considerations
would apply equally to any other Ansible playbook.

Recommendations
---------------

- Use SSH public key-based authentication to access target instances.

- Allow the SSH user to execute sudo commands without a password.

- Restrict access by time, rather than by command.

TPA needs access only when you are first setting up your cluster or
running ``tpaexec deploy`` again to make configuration changes, e.g.,
during a maintenance window. Until then, you can disable its access
entirely (a one-line change for both ssh and sudo).

During deployment, everything Ansible does is generally predictable
based on what the playbooks are doing and what parameters you provide,
and each action is visible in the system logs on the target instances,
as well as the Ansible log on the machine where tpaexec itself runs.

Ansible’s focus is less to impose fine-grained restrictions on what
actions may be executed and more to provide visibility into what it does
as it executes, so elevated privileges are better assigned and managed
by time rather than by scope.

SSH and sudo passwords
----------------------

We *strongly* recommend setting up password-less SSH key authentication
and password-less sudo access, but it is possible to use passwords too.

If you set ``ANSIBLE_ASK_PASS=yes`` and ``ANSIBLE_BECOME_ASK_PASS=yes``
in your environment before running tpaexec, Ansible will prompt you to
enter a login password and a sudo password for the remote servers. It
will then negotiate the login/sudo password prompt on the remote server
and send the password you specify (which will make your playbooks take
noticeably longer to run).

We do not recommend this mode of operation because we feel it is a more
effective security control to completely disable access through a
particular account when not needed than to use a combination of
passwords to restrict access. Using public key authentication for ssh
provides an effective control over who can access the server, and it’s
easier to protect a single private key per authorised user than it is to
protect a shared password or multiple shared passwords. Also, if you
limit access at the ssh/sudo level to when it is required, the passwords
do not add any extra security during your maintenance window.

sudo options
------------

To use Ansible with sudo, you must not set ``requiretty`` in
sudoers.conf.

If needed, you can change the sudo options that Ansible uses
(``-H -S -n`` ) by setting ``become_flags`` in the
``[privilege_escalation]`` section of ansible.cfg, or
``ANSIBLE_BECOME_FLAGS`` in the environment, or ``ansible_become_flags``
in the inventory. All three methods are equivalent, but please change
the sudo options only if there is a specific need to do so. The defaults
were chosen for good reasons. For example, removing ``-S -n`` will cause
tasks to timeout if password-less sudo is incorrectly configured.

Managing privilege escalation configuration
-------------------------------------------

Default sudo configuration
^^^^^^^^^^^^^^^^^^^^^^^^^^

By default, TPA automatically manages sudo-related configuration on
target instances, including installing the sudo package if not present
and configuring sudoers files for various components.

The default value of ``privilege_escalation_command`` is ``"sudo"`` ,
which enables TPA to manage sudo installation and configuration.

Using an alternative privilege escalation command
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

If your environment uses a different privilege escalation command, you
can configure TPA to use an alternative by setting
``privilege_escalation_command`` in ``cluster_vars`` :

.. code:: yaml

   cluster_vars:
     privilege_escalation_command: other_tool  # replaces sudo; may include arguments

The value is used as a direct in-place replacement for ``sudo`` in each
privilege escalation command TPA generates for managed applications. It
may include additional arguments if needed (e.g., ``other_tool --flag``
).

You can set it to any privilege escalation command supported by
Ansible’s become mechanism. Refer to the `Ansible privilege escalation documentation <https://docs.ansible.com/ansible/latest/playbook_guide/playbooks_privilege_escalation.html>`_ 

for the complete list of supported methods.

**Important:** ``sudo`` is the only privilege escalation command
officially supported by EDB. When using alternative commands, you are
responsible for ensuring compatibility and proper configuration. EDB
Support may have limited ability to assist with issues related to
alternative privilege escalation mechanisms.

Ansible’s become method vs. ``privilege_escalation_command``
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

There are two separate privilege escalation settings to be aware of:

- **``ansible_become_method``** controls how Ansible itself escalates
  privileges when running deployment tasks on target instances. This is
  an Ansible variable set in your inventory or ``ansible.cfg`` .

- **``privilege_escalation_command``** controls the command that the
  managed applications (EFM, repmgr, HARP) invoke at runtime to escalate
  privileges for service management operations, independently of
  Ansible.

Because they serve different purposes, both must be configured
consistently. If your environment uses an alternative tool instead of
sudo, you must set ``privilege_escalation_command`` in ``cluster_vars``
*and* configure ``ansible_become_method`` in your Ansible inventory or
``ansible.cfg`` to match.

Recommended approach: Using hooks
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

If you need to use an alternative privilege escalation command, we
recommend using :ref:`TPA hooks <TPA hooks>`  to configure your privilege escalation
mechanism. Hooks allow you to run custom tasks at specific points during
deployment, giving you full control over how privilege escalation is
configured whilst keeping your customisations separate from TPA’s core
deployment logic.

For example, you can use a ``post-repo`` hook to install and configure
your privilege escalation command after repositories are configured, or
a ``pre-deploy`` hook to set up the necessary permissions before the
main deployment begins. This approach provides better maintainability
and makes it easier to manage environment-specific requirements.

Manual configuration requirements
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

When using an alternative privilege escalation command (anything other
than ``"sudo"`` ), TPA will skip sudo package installation and sudoers
configuration. You must install and configure the chosen privilege
escalation command on all target systems, either manually or with hooks,
before running ``tpaexec deploy`` .

**1. Service management permissions for the postgres user**

Your privilege escalation mechanism must allow the postgres system user
to execute systemctl commands for starting, stopping, restarting, and
reloading PostgreSQL and related services. This is required for failover
managers (repmgr, HARP, EFM) to function correctly during automatic
failover operations.

For example, with sudo, TPA would configure the following permissions:

::

   postgres ALL=(ALL) NOPASSWD: /bin/systemctl start postgresql
   postgres ALL=(ALL) NOPASSWD: /bin/systemctl stop postgresql
   postgres ALL=(ALL) NOPASSWD: /bin/systemctl restart postgresql
   postgres ALL=(ALL) NOPASSWD: /bin/systemctl reload postgresql

You must configure equivalent permissions in your chosen privilege
escalation system.

**2. EFM database function permissions (EFM clusters only)**

If your cluster uses EFM as the failover manager, your privilege
escalation mechanism must allow the EFM system user to execute the
``efm_db_functions`` script as the postgres user. This is required for
EFM to perform health checks and failover operations.

For example, with sudo, TPA would configure:

::

   efm ALL=(postgres) NOPASSWD: /usr/edb/efm-X.Y/bin/efm_db_functions

Configure equivalent permissions in your privilege escalation system to
allow the efm user to run this script as the postgres user.

Logging
-------

For playbook executions, the sudo logs will show mostly invocations of
Python (just as it will show only an invocation of bash when someone
uses ``sudo -i`` ).

For more detail, the syslog will show the exact arguments to each module
invocation on the target instance. For a higher-level view of why that
module was invoked, the ansible.log on the controller shows what that
task was trying to do, and the result.

If you want even more detail, or an independent source of audit data,
you can run auditd on the server and use the SELinux log files. You can
get still more fine-grained syscall-level information from bpftrace/bcc
(e.g., opensnoop shows every file opened on the system, and execsnoop
shows every process executed on the system). You can do any or all of
these things, depending on your needs, with the obvious caveat of
increasing overhead with increased logging.

Local privileges
----------------

The :ref:`TPA installation <TPA installation>` 

mention sudo only as shorthand for “run these commands as root somehow”.
Once TPA is installed and you have run ``tpaexec setup`` , TPA itself
does not require elevated privileges on the local machine. (But if you
use Docker, you must run tpaexec as a user that belongs to a group that
is permitted to connect to the Docker daemon.)
